The short answer
An AI detector false positive is human writing that a detector labels as AI-written. It happens because detectors measure how predictable a text is, and people write predictable text too. Short texts, formulaic essays and writing by non-native English speakers are at higher risk. A score is a probability, not proof.
The main points of this guide:
- Every detector can be wrong. Vendors say so in their own documentation.
- Some writers are flagged more than others. Published research found a large gap between native and non-native English writers.
- Length matters. A short text gives a detector very little to measure.
- A score should start a conversation. It should never be the only basis for an academic integrity decision.
Why human writing gets flagged
A detector does not know who wrote a text. It estimates how likely each word is, given the words before it. AI text tends to pick likely words, so very predictable text looks like AI. Our guide on how AI detectors work explains the signals. The problem is that honest writing is often predictable as well.
| Kind of writing | Why it can look like AI |
|---|---|
| Short texts | There is too little text to measure. Turnitin wrote in 2023 that there “may not be enough signal” in submissions under 300 words. |
| Formulaic or template-like writing | A five-paragraph essay, a lab report or a cover letter follows a pattern. Patterns are predictable by design. |
| Generic introductions and conclusions | Turnitin’s release notes report more false positives in the first and last few sentences of a document, which are often “written in a generic way”. |
| Writing by non-native English speakers | A smaller range of words and sentence shapes makes text more predictable. A 2023 study measured this directly. |
| Lists, tables and other non-prose | Detectors are built for paragraphs. Turnitin says its model does not reliably handle bullet points, tables, annotated bibliographies, poetry, scripts or code. |
| Human sentences next to AI sentences | Turnitin says sentence-level mistakes are more common in documents that mix human and AI writing, especially where one changes to the other. |
Heavily edited text belongs in the formulaic group. Careful editing removes odd word choices and uneven sentences, and those are part of what makes writing look human to a detector. We found no vendor figure for how often this causes a flag, so we do not give one.
None of this means you should write differently to please a detector. Clear, well-organised writing is good writing. The answer to a false flag is evidence of how you wrote, not a different style.
See an AI check sentence by sentence
What the research says
The best-known study is “GPT detectors are biased against non-native English writers” by Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou of Stanford University, published in the journal Patterns in 2023. The authors ran seven widely used detectors on two sets of essays written by people.
| Essays tested | What the seven detectors did |
|---|---|
| 88 essays by US eighth-grade students | “Near-perfect accuracy”, in the authors’ words. |
| 91 TOEFL essays by non-native English writers | Wrongly flagged more than half as AI-generated. The average false positive rate was 61.22%. |
| The same 91 essays, flagged by all seven detectors | 18 essays, or 19.78%. |
| The same 91 essays, flagged by at least one detector | 89 essays, or 97.80%. |
The authors also looked at why. The essays that every detector flagged had significantly lower perplexity than the rest, which means their wording was more predictable. The paper’s explanation is that non-native writers tend to use a narrower range of words and structures, and detectors read that as a sign of AI. The authors “caution against” using these detectors “in evaluative or educational settings”, especially where non-native English speakers could be penalised.
The limits of that study
It is fair to say what the study does not show. The authors call it a pilot study and say the samples are relatively small. It was done in 2023, and detectors have changed since. Turnitin’s detector was not one of the seven. Turnitin has also pointed out that the TOEFL essays were all under 150 words long, which is shorter than its tool will accept.
So the study does not give a false positive rate for every detector today. It does show that the risk is not spread evenly. Writers with a smaller English vocabulary carry more of it.
What Turnitin says about its own tool
Many students meet AI detection through Turnitin, so its documents are worth reading closely. These are Turnitin’s own statements about its own product. They are not independent tests.
- It can be wrong. The guide Using the AI Writing Report says the model “may not always be accurate (it may misidentify human-written, AI-generated, and AI-paraphrased text), so it should not be used as the sole basis for adverse actions against a student”.
- Low scores are hidden. The same guide says there is “a higher incidence of false positives” when the percentage is between 0 and 19. Scores above 0% and below 20% are shown as an asterisk, with no number and no highlights.
- Short texts are not scored. A file needs at least 300 words of prose in a long-form format. Turnitin’s model release notes say a score on less than 300 words “is likely less accurate”.
- Its stated document-level rate. In a June 2023 blog post, Turnitin gave a document false positive rate of less than 1% for documents with 20% or more AI writing.
- Its stated sentence-level rate. The same post gave a sentence-level false positive rate of around 4%, and said that 54% of the time those sentences sit right next to actual AI writing.
- Non-native English writers. In an October 2023 post, Turnitin reported its own test on essays by English language learners. For documents of 300 words or more it found a small difference between the two groups that was not statistically significant. For shorter documents the difference was larger, and the rate was significantly above its 1% target.
Put together, the vendor’s message matches the research on one point: short, simple texts are where detection is weakest. We explain the report itself in does Turnitin detect AI?
A note on platforms. An AI score does not come from the learning management system. It comes from the tool a school has connected to it. See does Blackboard detect AI?, does Canvas have AI detection? and does SafeAssign detect AI?
How common are false positives?
There is no single honest number. The rate changes with three things:
- The tool. Each detector is a different model with a different cut-off. The same essay can get different scores from different tools.
- The text. Length, genre and how formulaic the writing is all move the result.
- The writer. The 2023 study above found near-perfect results for one group and a false positive rate over 60% for another, with the same detectors.
Scale matters too. A small rate still means real people. If 1 in 100 honest essays is flagged and an instructor grades 200 essays in a term, about two honest students are flagged. That is simple arithmetic, not a measured figure, but it is the reason a flag has to be checked before anyone acts on it.
If you are falsely flagged
You cannot argue a score down. You can show how the work was written. Keep it calm and factual.
- Ask to see the report. Ask which tool produced it and which passages were flagged. Two sentences in a conclusion and a whole essay are very different situations.
- Gather your drafts and version history. A document that grew over several days, with edits and deletions, is the best evidence you have. Share the original file, not a copy.
- Collect your notes and sources. Outlines, reading notes and earlier drafts show your thinking.
- Talk to your instructor. Offer to explain your argument and how you wrote the piece. Being able to discuss your own work is strong evidence.
- Leave the submitted text alone. Do not rewrite honest work to change a number, and do not edit the file or its history. Changing it afterwards makes your evidence weaker.
- Read your school’s procedure, including how to appeal.
That is the short version. The full checklist, with what to ask for and what not to do, is in how to prove you didn’t use AI.
If English is not your first language, say so. It is relevant. Published research found that detectors flag non-native writers more often, and it is reasonable to ask that this is taken into account.
Before a teacher acts on a score
A flag is a reason to look closer. Turnitin’s March 2023 post on false positives says the company “does not make a determination of misconduct” and that instructors need to apply their professional judgment, their knowledge of their students and the context of the assignment.
- Have a process before you need one. Turnitin’s advice is to consider the possibility of a false positive up front and to tell students what your approach will be.
- Check whether the text suits the tool. Is it long enough? Is it prose? A short answer, a list or a reflection of a few lines is where detection is weakest.
- Think about who wrote it. A student writing in a second language is more likely to be flagged by mistake.
- Look at where the flags fall. A flagged introduction or conclusion means less than a flagged body.
- Compare with what you know. Earlier work, in-class writing and drafts tell you more than a percentage.
- Talk first. Turnitin’s June 2023 post says to use highlighted sentences “to initiate a conversation, not to draw a conclusion”. Ask the student to explain the work and the process they followed.
- Give the benefit of the doubt. The March 2023 post says that if the evidence is unclear, assume students will act with integrity.
- Share the report with the student, and follow your institution’s procedure.
For choosing a threshold, see what percentage of AI is acceptable?
How to read an AI score
An AI score is a probability, not proof. It tells you how much a text resembles AI writing to one model. It does not tell you what happened when the text was written.
The Plagiarism Checker Plus AI detector works the same way and has the same kind of limits. It gives an overall AI score and a likelihood for every sentence, so you can see which passages drive the number instead of guessing from one figure. It can flag human writing, like any detector. It should never be the only basis for an academic integrity decision.
Used well, a sentence-level report helps in two ways:
- For a student, it shows which parts of a draft a reader might question, so you know which notes and drafts to keep ready. You do not need to rewrite honest work.
- For a teacher, it shows where to look and what to ask about, before any conversation with the student.