Guidelines
A primary goal of our guidelines is to enable reproducibility and replicability of empirical SE studies involving LLMs. As repeating LLM-focused research to verify results lies somewhere between the ACM’s definitions of reproducibility (different team, same research artifacts) and replicability (different team, different research artifacts) (ACM Publications Board 2020) due to potential model changes and incomplete research artifacts (e.g., unreported prompts, configurations, or seeds), we follow Angermeir et al. (2025) and use the terms interchangeably. While previous guidelines regarding open science and empirical studies still apply, LLM-specific characteristics (e.g. inherent non-determinism (Song et al. 2025; Yuan et al. 2025), opaque and proprietary models) present additional replicability challenges, which, in turn, demand new guidance.
Each guideline below begins with a brief summary, followed by its rationale, recommendations, examples, benefits, and challenges, with links to the relevant study types. The rationale articulates the underlying principle, i.e., why the guideline matters, while the recommendations provide concrete, actionable practices.
The guidelines further contain advice for reviewers and close with a see also list linking related guidelines. Broadly, reviewers should use our guidelines to help them interrogate the extent to which a manuscript’s authors have done what was practically possible to improve reproducibility. We must neither accept research absent reasonable efforts to improve reproducibility, nor reject research for failing to obtain an impossible goal. We borrow this principle of “reasonable efforts” from the SIGSOFT Empirical Standards (Ralph et al. 2021), where it applies to methodological rigor more generally. Of course, papers should acknowledge their limitations, but to determine whether these limitations are reasonable, reviewers should ask “have the authors done what they could to minimize limitations?” Our guidelines attempt to capture what “reasonable efforts” practically means for LLM-based studies.
To distinguish essential criteria from recommendations, our guidelines use two tiers. A must criterion is a requirement. Studies that intend to follow these guidelines are expected to meet all must criteria. A should criterion represents a desired practice that strengthens a study’s rigor or transparency. However, there may be valid reasons to deviate in particular circumstances (e.g., resource constraints, inapplicability to a specific study context or type). Nonetheless, authors should briefly justify any deviation from a should criterion that could compromise a study’s validity or reproducibility. The distinction between must and should also reflects what the criterion mandates, not only its severity. A must criterion is typically a disclosure obligation such as reporting model versions, publishing prompts, describing architectures, and discussing limitations, so that readers can independently evaluate the choices authors made. A should criterion is typically a methodological recommendation such as which baselines to include and which validation strategies or statistical analyses to apply.
The following sections indicate which information we expect researchers to report, and whether it should be in the paper or supplementary material. Where a publication venue’s page limits hinder reporting all expected elements in the paper, it is better to report essential information in the supplementary material than not at all. The supplementary material should be published according to the ACM SIGSOFT Open Science Policies (Graziotin 2024).
For a compact overview of each guideline’s rationale and core recommendations, see the summary. For an item-by-item view to walk through during reporting, see the reporting checklist.
Guidelines by Study Type
Note: Each guideline’s study-type-specific guidance is detailed in the corresponding subsection.
The guideline’s core recommendations:
● = must be followed for this study type.
● = should be followed for this study type.
– = are not directly applicable for this study type.
| Annotators | Judges | Synthesis | Subjects | Usage | Tools | Benchmarking | |
|---|---|---|---|---|---|---|---|
| Declare Usage | ● | ● | ● | ● | ● | ● | ● |
| Model Version | ● | ● | ● | ● | ● | ● | ● |
| Design | ● | ● | ● | ● | ● | ● | ● |
| Traces | ● | ● | ● | ● | ● | ● | ● |
| Benchmarks & Metrics | ● | ● | ● | ● | ● | ● | ● |
| Open LLM | ● | ● | ● | – | – | ● | ● |
| Human Validation | ● | ● | ● | ● | ● | ● | ● |
| Limitations | ● | ● | ● | ● | ● | ● | ● |
References
ACM Publications Board. 2020. “Artifact Review and Badging – Current.” https://www.acm.org/publications/policies/artifact-review-and-badging-current.
Angermeir, Florian, Maximilian Amougou, Mark Kreitz, Andreas Bauer, Matthias Linhuber, Davide Fucci, Fabiola Moyón Constante, Daniel Méndez, and Tony Gorschek. 2025. “Reflections on the Reproducibility of Commercial LLM Performance in Empirical Software Engineering Studies.” CoRR abs/2510.25506. https://doi.org/10.48550/ARXIV.2510.25506.
Graziotin, Daniel. 2024. “Acmsigsoft/Open-Science-Policies: V1.0.0.” Zenodo. https://doi.org/10.5281/zenodo.10796477.
Ralph, Paul, Nauman bin Ali, Sebastian Baltes, Domenico Bianculli, Jessica Diaz, Yvonne Dittrich, Neil Ernst, et al. 2021. “Empirical Standards for Software Engineering Research.” https://arxiv.org/abs/2010.03525.
Song, Yifan, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. 2025. “The Good, the Bad, and the Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism.” In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, edited by Luis Chiruzzo, Alan Ritter, and Lu Wang, 4195–4206. Association for Computational Linguistics. https://doi.org/10.18653/V1/2025.NAACL-LONG.211.
Yuan, Jiayi, Hao Li, Xinheng Ding, Wenya Xie, Yu-Jhe Li, Wentian Zhao, Kun Wan, Jing Shi, Xia Hu, and Zirui Liu. 2025. “Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference.” In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025. https://openreview.net/forum?id=Q3qAsZAEZw.