The short answer
AI detectors work as a signal, not as proof. They are right more often than chance, and they do best on long text that came straight from an AI tool. They also make mistakes in both directions: they flag some human writing and they miss some AI writing. A score is a probability, never a verdict.
- They do measure something real. AI text has patterns that software can pick up.
- They are not reliable enough to decide a case. Published studies found errors with every tool they tested.
- The errors are not evenly spread. Short text, edited text and some kinds of human writing are misjudged more often.
What a detector measures
A detector does not find a record of someone using an AI tool. It reads the text and estimates how likely it is that a language model produced it. The main clue is predictability: AI tools tend to pick likely words, so their text is smoother and more even than most human writing.
That is why the result is a likelihood. The detector is saying “this text looks like AI text”, not “an AI wrote this”. Our guide on how AI detectors work explains the signals in detail, so we will not repeat them here.
See what an AI check shows for your text
Where detectors do well
Detectors are most reliable in two situations.
- Recognising ordinary human writing. In the Weber-Wulff study of 14 detection tools, overall accuracy on documents written by people was 96%.
- Long text taken straight from an AI tool. The same study measured 74% accuracy on AI-generated documents that nobody had changed. That is well above a coin toss, and well below certain.
Length helps because it gives the detector more to measure. Turnitin describes the other side of this in its AI writing detection FAQ: in documents of only a few hundred words, its prediction is mostly “all or nothing”.
Where detectors fail
Four situations cause most of the mistakes.
1. Short text
A paragraph or a short answer has too few words for a steady estimate. Turnitin says that in a short document, text that is a mix of AI-generated and original content could be flagged as entirely AI-generated.
2. Edited and mixed text
Real documents are often a mix: some AI text, some human text, and AI text that a person has reworked. Detectors handle this badly. In the Weber-Wulff study, accuracy fell from 74% on unchanged AI text to 42% on AI text that a person had edited afterwards. Turnitin also says that in a longer document with both kinds of writing it can be difficult to tell exactly where the AI writing begins and the original writing ends.
This is a limit of the tools. It is not a way around a course rule. If AI writing is not allowed in your course, editing it does not make it allowed.
3. Formulaic human writing
Some human writing is very regular. Turnitin’s FAQ says its false positives can include content without a lot of structural variation, text that literally repeats itself, and text that has been paraphrased without developing new ideas. Lab reports, set-format answers and template-style essays can all read this way.
4. Writing by non-native English speakers
This is the best documented problem. Liang and colleagues ran seven widely used detectors on 91 English test essays written by non-native speakers and on 88 essays by US eighth-grade students. The detectors classified the US essays accurately. They wrongly labelled more than half of the non-native essays as AI-generated, with an average false-positive rate of 61.3%.
The authors explain why. Writers working in a second language tend to use a smaller range of words and sentence shapes. That makes the text more predictable, and predictable text is exactly what a detector treats as a sign of AI. Our guide to AI detector false positives covers what this means for students and teachers.
What the research found
Two peer-reviewed studies are cited more than any others. Here is what each one tested and found.
| Study | What was tested | Main finding |
|---|---|---|
| Liang et al., 2023 (Patterns) | Seven detectors on 91 essays by non-native English writers and 88 essays by US eighth graders | The US essays were classified accurately. The non-native essays had an average false-positive rate of 61.3%. All seven detectors agreed that 19.8% of them were AI-written, and they were not. |
| Weber-Wulff et al., 2023 (International Journal for Educational Integrity) | 12 publicly available tools and two commercial systems, on human-written, machine-translated, AI-generated and edited AI documents | Every tool scored below 80% accuracy and only five scored above 70%. The tools leaned towards calling text human-written. Accuracy was 96% on human text, 74% on unchanged AI text and 42% on AI text edited by a person. |
The two studies point at different errors. Liang shows that human writing can be flagged, and that it happens more to some writers than to others. Weber-Wulff shows that AI writing is often missed. The Weber-Wulff authors conclude that the tools they tested are unsuitable as evidence of academic misconduct.
Vendors publish their own figures as well. Turnitin, for example, states a false-positive rate of less than 1% for its detector. That is a company’s measurement of its own product, and we have not tested it. Our guide on whether Turnitin detects AI goes through what Turnitin says and where it sets its own limits.
Can AI detectors be wrong?
Yes. There are two kinds of mistake, and they hurt different people.
| Mistake | What happens | Who it affects |
|---|---|---|
| False positive | Human writing is flagged as AI | A writer who did the work and now has to defend it |
| False negative | AI writing is passed as human | Everyone who trusted a clean result |
A tool can lower one kind of mistake only by accepting more of the other. Turnitin says this openly: to keep its false-positive rate low, there is a chance it might miss some AI-written text in a document. So a low score does not prove a text is human, and a high score does not prove it is AI.
Two detectors can also disagree about the same text, because each one is built and trained differently. A result from one tool does not tell you what another tool will show.
How to use a score responsibly
An AI score should never be the only basis for an academic integrity decision. Turnitin gives the same warning about its own tool: the percentage “should not be used as the sole basis for action or a definitive grading measure by instructors”. In practice that means:
- Treat the score as a reason to look closer. It tells you where to read carefully. It does not tell you what happened.
- Read the flagged sentences, not just the total. A few flagged lines in a long document mean something different from whole flagged paragraphs.
- Ask whether the text is a known weak spot. Short, formulaic or second-language writing is misjudged more often, so give the score less weight there.
- Look for other evidence. Drafts, notes, version history and the writer’s earlier work show how a text was produced. A detector cannot.
- Talk to the writer. Someone who wrote a piece can usually explain it.
- Follow the written policy. Your institution’s process decides what counts as evidence and how a person can respond.
If you are the one who was flagged, see how to prove you didn’t use AI. If you are wondering whether a certain number is safe, see what percentage of AI is acceptable? For how teachers put a score together with other evidence, see can professors detect ChatGPT?
Checking your own writing
The Plagiarism Checker Plus AI detector gives an overall AI score and a likelihood for every sentence, so you can see which lines the result is based on. Students checking an essay can start from the AI detector for essays.
Everything on this page applies to our tool as well. Its score is a probability, not proof. It can be wrong about human writing and about AI writing. It is a different tool from the one your school or client uses, so the two can disagree. Use it to find passages worth rereading, and keep your drafts.