Research

Study of 135,389 documents: AI detectors disagree widely on non-native writers' human text

Researchers at Chung-Ang University in South Korea and the editing company Wordvice have tested 13 publicly available AI text detectors on 135,389 documents written by non-native English speakers. Their paper, posted to arXiv on 27 August 2026, found that the detectors gave very different verdicts on the same human-written texts, and that professional human editing pushed some detectors' scores up and others' down. The study did not test commercial detectors, and it has limits that matter for anyone applying it to student work. Here is what it measured.

Key figures from an August 2026 study of AI detectors: 135,389 document pairs by non-native English writers, 13 detectors tested, and false positive rates on pre-ChatGPT human text ranging from 0.0% to 100.0%.

Summary

  • A paper posted to arXiv on 27 August 2026 tested 13 publicly available AI detectors on 135,389 documents by non-native English writers, each paired with a version edited by a human native-speaker editor.
  • On documents from 2018 to 2022, before ChatGPT, the share wrongly flagged as AI ranged from 0.0% to 100.0% depending on the detector.
  • Human editing lowered AI scores on some detectors and raised them on others. The authors say the study does not cover proprietary commercial detectors and may not apply to student essays.

Key takeaways

  • Different detectors can reach opposite conclusions about the same human-written text, so a single score says a lot about the tool as well as the writing.
  • Getting human help to polish English changed detector results in this study, even though no AI wrote the text.
  • The authors conclude that detector output should not be treated as definitive proof of AI use, particularly where editing is common.

What the study did

The paper is Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing by Hyeonchu Park, Gahye Jeong and Bugeun Kim. It was submitted to arXiv on 27 August 2026, which makes it about five weeks old at the time of writing.

Earlier research, which the authors cite, reported that detectors flag non-native English writing more often than native writing. The authors argue that those comparisons are hard to interpret, because two groups of writers also differ in topic, subject and purpose. Their approach was to compare each document with itself.

They used records from Wordvice, a professional English editing service. Each record has two versions of one document: the original by a non-native English speaker, and the version edited by a native-speaker editor. According to the paper, the company's rules allow editors to use AI grammar checkers only for mechanical errors and prohibit AI rewriting or paraphrasing.

  • Sample: 135,389 document pairs from 2018 to 2025, across more than 40 subject areas.
  • Before ChatGPT: 104,752 pairs (77.4%) date from 2018 to 2022. The other 30,637 date from 2023 to 2025.
  • Type of document: 52.9% came through academic editing and 37.6% through admission editing. Application essays were the largest single subject group, with 22,639 documents.
  • Detectors: 13 publicly available tools of three kinds (token statistics, zero-shot and trained classifiers), run as published with no tuning.
  • Decision rule: each score was converted to a 0 to 1 scale and anything above 0.5 counted as flagged.

What it found

The first test used only the 2018 to 2022 documents. The authors treat these as human-written with negligible AI involvement, so any flag counts as a false positive. The abstract reports that rates varied "from 0.0% to 100.0%" across the 13 detectors.

Detector (type)Original flaggedEdited flagged
Log-rank (token statistics)100.0%100.0%
GLTR (token statistics)99.9%99.9%
RoBERTa (classifier)93.1%93.9%
RADAR (classifier)88.3%86.4%
Fast-DetectGPT (zero-shot)25.2%31.6%
LastDE+ (zero-shot)19.9%27.3%
MAGE (classifier)16.7%7.6%
BiScope (zero-shot)0.6%0.3%
Binoculars (zero-shot)0.0%0.1%

The figures above are from Table 3 of the paper, which it labels as the 2018 to 2022 baseline. Nine of the 13 detectors are shown. The four detectors based on token statistics flagged almost every document. Several zero-shot detectors flagged almost none.

The second finding is about editing. Across the full dataset, the paper reports that human editing cut MAGE's false positive rate by 10.7 percentage points and RADAR's by 4.7, while raising Fast-DetectGPT's by 4.6 and LastDE+'s by 5.8. The edits were the same. The detectors read them in opposite ways.

The size of the shift tracked the amount of editing: heavier edits produced larger score changes in the same direction for each detector. The authors add that the shifts were large enough to matter in practice for only 7 of the 13 detectors, and that the editing effect was weaker in the 2023 to 2025 documents.

What the study does not show

The paper lists its own limitations, and they narrow what can be claimed.

  • No commercial detectors. The authors state that the analysis does not cover proprietary commercial systems. It says nothing direct about the products universities buy.
  • No native-speaker comparison. Every original was written by a non-native English speaker, so the study does not measure a gap between native and non-native writers.
  • Not classroom essays. The authors say detector behaviour may differ for student essays and other kinds of writing.
  • One threshold. Absolute rates depend on the 0.5 cut-off and on how each score was rescaled. The authors say the before-and-after comparisons are more robust than the absolute rates.
  • No AI-written controls. The study cannot say whether editing affects AI-generated text in the same way.
  • Correlation only. The links between particular language features and score changes are not shown to be causal.

Readers should also know that one of the three authors is affiliated with Wordvice, the editing company that supplied the data. The paper's acknowledgements list funding from South Korean government research programmes.

What this means for students and teachers

The practical points below follow from the reported findings. They are our reading, not the authors' wording, apart from the one quoted conclusion.

For students writing in a second language

  • Keep your drafts, notes and version history. In this study, human-written text was flagged by some detectors and cleared by others, so a record of how you wrote is stronger evidence than any score.
  • If a proofreader or editing service worked on your text, keep the original and the edited version. Check your institution's rules on proofreading first.
  • If you are flagged, ask which tool was used and what threshold was applied. The study shows that both change the result.

For teachers and integrity officers

  • Treat a detector score as a reason to look more closely, not as a finding. The authors write that detector outputs "should not be treated as definitive evidence of AI use".
  • Ask your vendor for false positive data on non-native and professionally edited writing. This study did not test commercial tools, so the question is open for them.
  • Polished English is not evidence either way. Some detectors in the study scored edited text as more human, others as more AI-like.

Sources

FAQ

Are AI detectors biased against non-native English speakers?

This study does not settle that, because it had no native-speaker comparison group. It found that 13 publicly available detectors gave false positive rates from 0.0% to 100.0% on pre-ChatGPT documents by non-native writers, so the answer depended heavily on which detector was used. It did not test commercial detectors.

Can human proofreading make my writing look AI-generated?

In this study it depended on the detector. After professional human editing, false positive rates rose on Fast-DetectGPT and LastDE+ and fell on MAGE and RADAR. The authors say the findings may not carry over to student essays.

Which AI detectors did the study test?

Thirteen publicly available research detectors: Log-rank, Log-likelihood, Entropy, GLTR, Fast-DetectGPT, LastDE+, DetectLLM-LRR, DiVeye, BiScope, Binoculars, RoBERTa, MAGE and RADAR. The authors state that proprietary commercial detectors were not covered.

Is an AI detector score enough to prove a student used AI?

The authors of this study say no. They conclude that detector outputs should not be treated as definitive evidence of AI use, especially in academic settings where editing is common, and should be weighed alongside contextual information.

What happens next

The arXiv listing says the paper was presented in EMNLP 2026, a natural language processing conference. The authors say the next step is to add matched AI-generated texts with controlled amounts of human editing, and to test other kinds of writing. Until work like that covers student essays and commercial detectors, this study is evidence about how detectors react to writing style, not a rating of any tool a university uses.

Next step

How AI detectors work

A plain-English guide to the methods behind AI-writing detectors and why their scores can differ.

Read the guide