The writer's problem
A journal editor wondering how much of the language in incoming peer reviews and papers is now AI-assisted cannot ask every reviewer directly. A Stanford team led by Weixin Liang and James Zou built a way to estimate it from word patterns across a whole corpus instead of judging any single document. Their paper on AI-modified peer reviews, accepted to the 2024 International Conference on Machine Learning, applies that method to a specific, named set of venues.
What the documents show
The paper's own abstract states its maximum-likelihood model, which tracks shifts in the frequency of specific words such as 'commendable' and 'meticulous' before and after ChatGPT's release, estimates that between 6.5 and 16.9 percent of text submitted as peer reviews to ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023 was substantially modified by an LLM. The estimated share was higher in reviews reporting lower confidence, submitted close to the deadline, or from reviewers less likely to respond to author rebuttals. A companion study by an overlapping author team applied the same word-frequency method to a much larger corpus. That paper analyzed 950,965 papers on arXiv, bioRxiv and in Nature portfolio journals published between January 2020 and February 2024, finding up to 17.5 percent estimated LLM modification in Computer Science papers and as little as 6.3 percent in Mathematics and the Nature portfolio, a considerably wider range across scientific writing generally than the peer-review-specific estimate.
The editorial choice
A journal should not treat either percentage as a rule for any individual submission; both papers state their method operates at the corpus level and is not designed to accuse a specific person's specific text. Editorially, the more defensible use of this research is to inform disclosure policy discussions at the level of a field or venue, using the peer-review figure for review text specifically and the broader figure only for its own, differently composed corpus of papers.
What stays with the author
Neither paper tells an editor or a reviewer what to do about legitimate language editing, including editing by non-native English speakers using AI tools for clarity, and both estimate prevalence without judging any individual case. Deciding what counts as an acceptable use of a drafting or editing tool in a review or a manuscript remains a policy choice for the journal, not a conclusion these measurements draw.
- Does a field-wide estimate change what a journal's own disclosure policy should require?
- Would the same method, applied to a different venue or year, find a different rate?
- How should a wide estimate range, 6.5 to 16.9 percent, be reported without implying false precision?
Both papers describe a scale of change across large corpora of scientific writing, not a verdict on any single review or paper's honesty.
Follow the source.
The paper's own abstract and results report a corpus-level, word-frequency-based estimate that 6.5 to 16.9 percent of peer-review text at ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023 was substantially LLM-modified.
Source date: 11 Mar 2024 · Retrieved: 16 Sept 2026
This related paper's own abstract reports the same word-frequency method applied to 950,965 papers across arXiv, bioRxiv and Nature portfolio journals, estimating up to 17.5 percent LLM modification in Computer Science and as little as 6.3 percent in Mathematics and the Nature portfolio.
Source date: 1 Apr 2024 · Retrieved: 16 Sept 2026
Site publication is not established by an event date. Original record ID: 0030-bf-098. This local design review does not change its editorial status.