One in five reviewers used AI despite a ban, ICML 2026 study finds
The International Conference on Machine Learning (ICML), one of the largest research conferences in its field, ran an experiment during its 2026 review process: some reviewers were told they could not use AI language models at all, and others were allowed limited use. A paper describing the results was posted to arXiv on 16 September 2026 by 11 researchers, including the conference's program chairs. It reports that the two rules produced almost identical decisions on papers, and that many reviewers did not follow whichever rule they were given. In an anonymous survey, 22.5% of reviewers under the ban said they had used a language model anyway.

Summary
- ICML 2026 assigned each reviewer to one of two policies: no use of large language models (LLMs), or limited use for understanding papers and polishing the reviewer's own text.
- Policy assignment had near-zero effects on final decisions, paper scores and reviewer confidence, the paper reports. Reviews under the permissive policy were 5.5% to 7% longer.
- In an anonymous survey, 165 of 732 reviewers under the ban (22.5%) reported using an LLM, and 155 of 425 under the permissive policy (36.5%) reported a use that policy did not allow.
Key takeaways
- This is a rare randomised test of an AI-use rule, run on a live process with real consequences.
- The conference did enforce the ban. In March 2026 it rejected 497 papers whose authors had broken it as reviewers. Non-compliance was still common.
- The authors conclude that rules about which AI uses are allowed are hard to enforce and that clear guidance and norms may work better than strict bans. That is their interpretation, drawn from one conference.
What ICML did
The paper is Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026 by Sunnie S. Y. Kim, Wesley Hanwen Deng and nine co-authors from Microsoft Research, the University of Pennsylvania, Google Research, the University of Wisconsin-Madison, EPFL and Carnegie Mellon University. It has been posted as a preprint, which means it has not necessarily been peer reviewed itself.
According to the paper, ICML 2026 received 24,661 submissions from 76,159 authors and had 17,886 reviewers. Each reviewer worked under one of two policies.
| Policy | What it allowed | Reported LLM use (survey) | Reported rule-breaking (survey) |
|---|---|---|---|
| Conservative | No LLM use at any stage of reviewing, apart from incidental use in tools such as search and spelling checkers | 22.5% | 22.5% (165 of 732) |
| Permissive | LLMs to help understand the paper and related work, and to polish the review. Not for evaluating the paper or drafting the review | 91.5% | 36.5% (155 of 425) |
Authors and reviewers stated a preference first. Authors of 29.4% of main-track papers required the conservative policy. Among reviewers who were asked, 16.2% preferred the conservative policy, 33.9% preferred the permissive one, and 49.9% said they were fine with either. The experiment then randomised within those preferences. For example, 10% of the 16,877 papers whose authors accepted either policy were switched at random to the conservative one.
What the experiment found
The headline result is that the rule did not change outcomes. Among papers randomly switched to the no-AI policy, 73.0% were rejected or withdrawn. Among those left under the permissive policy the figure was 73.5%. Acceptance as a regular paper was 24.4% in both groups. Paper scores and reviewer confidence were also nearly identical.
There were small differences. Reviews written under the permissive policy were 5.5% to 7% longer. In one of the two analyses they were also rated slightly higher for quality by the area chairs who oversee reviewers.
The authors also ran a sample of reviews through Pangram, a commercial classifier of AI-generated text. It classed 52.2% of sampled reviews under the conservative policy as human-written, compared with 37.0% under the permissive policy. The paper presents this as a supplementary analysis and says compliance rates are uncertain because the available ways of measuring LLM use are limited. A classifier result is an estimate, and the paper does not treat it as proof about any individual review.
How the ban was enforced
ICML did not rely on general AI-text detectors to enforce its ban. In a statement on 18 March 2026, the program chairs described a different method. Each paper sent to a no-AI reviewer was a uniquely modified PDF with hidden instructions. If the file was uploaded to a language model, the instructions told the model to include two specific phrases, chosen at random from a dictionary of about 170,000, in its output. Reviews were screened for those phrases and, the chairs said, every flagged case was checked by a person.
| Date (2026) | What happened | Source |
|---|---|---|
| 28 January | Full-paper submission deadline. The experiment period begins. | arXiv paper |
| 18 March | ICML says 795 reviews by 506 no-AI reviewers were found to have used LLMs. It desk-rejects 497 papers submitted by 398 of those reviewers, about 2% of all submissions. | ICML blog |
| 30 April | Authors are notified of decisions. The experiment period ends. | arXiv paper |
| 6 to 11 July | The conference takes place in Seoul. | arXiv paper |
| 16 September | The experiment and survey results are posted to arXiv. | arXiv paper |
The chairs acknowledged that the method would miss cases. A reviewer could find and remove the hidden text, or the model could ignore it. In tests before the deadline, they said, success rates were over 80% for most models. The research-integrity news site CASRAI described the rejections as one of the most concrete enforcement actions taken so far against undisclosed AI use in peer review.
Put together, the two sources show a gap. The hidden-phrase check flagged around 1.5% of reviews under the conservative policy, according to the paper. The anonymous survey, taken after the penalties were public, put self-reported use at 22.5%. The check only catches one behaviour, uploading the PDF to a model that then follows the hidden instruction, so the two numbers measure different things.
Why reviewers said they used AI
The paper says non-compliant use was often driven by workload and by uncertainty about where the line was. Submissions to ICML rose from 6,538 in 2023 to 24,661 in 2026, an increase of 277%. The most common uses reported across all reviewers were editing or polishing their own text (73.8%), understanding the paper (66.5%) and reviewing related work (37.5%). Among permissive-policy reviewers, the prohibited uses reported included brainstorming feedback or questions (24.2%), drafting an outline or parts of the review (16.2%) and identifying a paper's strengths or weaknesses (14.8%).
Survey respondents broadly supported AI for assistive tasks and were more cautious about tasks that need a reviewer's judgement. Many doubted that such a boundary could be enforced.
A related account appeared the same week. Times Higher Education reported on 17 September that Nihar Shah, a co-author of the ICML paper and editor-in-chief of the journal Transactions on Machine Learning Research, had interviewed seven authors of desk-rejected papers and found that three could not answer basic questions about their own submissions. THE reported his figure that the journal's desk-rejection rate had risen from 6% in 2023 to 53%.
What this means for students and teachers
Peer reviewers are researchers, not students, and a conference is not a classroom. The study is still relevant to anyone writing an AI rule for a course, because it tested the two options most syllabuses choose between: a full ban, and permission for limited help.
- Teachers: a rule that separates allowed AI help (polishing) from disallowed help (drafting) was broken by more than a third of the surveyed reviewers who worked under it. If you set a rule like this, give worked examples of each side of the line. The paper's authors recommend the same for conferences.
- Teachers: enforcement here rested on a specific technical check with human verification, and it still caught only a small share of what reviewers later admitted to. The authors suggest building norms and guidance where a restriction cannot be enforced.
- Students: expert adults under time pressure gave the same reasons students give: too much work and unclear boundaries. If a course rule is unclear, ask before the deadline and keep a note of the answer.
- Research students: check the reviewing policy of any venue you review for. At ICML 2026 a breach led to the reviewer's own paper being rejected.
Sources
- Sunnie S. Y. Kim, Wesley Hanwen Deng, Jennifer Wortman Vaughan, Buxin Su, Weijie Su, Alekh Agarwal, Sharon Li, Martin Jaggi, Daniel G. Goldstein, Nihar B. Shah and Miroslav Dudík, Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026, arXiv:2609.19420, 16 September 2026.
- ICML 2026 Program Chairs, On Violations of LLM Review Policies, ICML Blog, 18 March 2026.
- CASRAI Editorial Board, ICML 2026 Desk-Rejects 497 Papers Over AI Reviews, CASRAI, 24 July 2026.
- Frances Jones, Researchers submitting AI-assisted papers they 'don't understand', Times Higher Education, 17 September 2026.
- Retraction Watch, Weekend reads, 26 September 2026 (lists the ICML paper among the week's research integrity items).
FAQ
Did ICML 2026 ban AI in peer review?
Partly. It ran two policies side by side. Under the conservative policy reviewers could not use language models at any stage. Under the permissive policy they could use them to understand a paper and related work and to polish their own review, but not to evaluate the paper or draft the review.
How many reviewers broke the no-AI rule?
Two figures exist and they measure different things. ICML's hidden-phrase check found 795 reviews by 506 no-AI reviewers that had used a language model. Separately, in an anonymous survey after the process, 165 of 732 no-AI reviewers who answered (22.5%) said they had used one.
Did allowing AI change which papers were accepted?
Not measurably. The paper reports near-zero effects of policy assignment on final decisions, scores and reviewer confidence. Reviews under the permissive policy were 5.5% to 7% longer.
How did ICML detect AI use by reviewers?
It did not use general AI-text detectors for enforcement. It hid instructions in each PDF that told a language model to insert two specific phrases into its output, then screened reviews for those phrases and had a person verify each flagged case.
What happened to reviewers who were caught?
According to ICML's March 2026 statement, 497 papers submitted by 398 of those reviewers were desk-rejected, which was about 2% of all submissions, and the flagged reviews were removed.
What happens next
The findings come from one conference in one review cycle, and the authors say both the tools and the norms around them are changing quickly. They suggest future work should pair surveys with direct measures of how reviewers use AI, for example through tools the conference itself provides. We will update this article if the paper is revised or if ICML announces its policy for 2027.
How AI detectors work
A plain-English guide to the methods behind AI-writing detectors, and what a result can and cannot tell you.


