The short answer
A plagiarism checker splits your text into short phrases and searches for those phrases in an index of web pages, publications and, in some tools, earlier student papers. Where your wording matches a source, it highlights the passage, names the source and adds the match to a percentage.
In short, a checker does four things:
- It breaks your text into small pieces.
- It looks for those pieces in the sources it has stored.
- It lines up your text with each source to find the matching passages.
- It reports the matches and a score.
What your text is compared with
A checker can only match your text against sources it has collected. That collection is called an index. It is built before you ever run a check, and it is different for every tool.
| Kind of source | What it is | Which checkers have it |
|---|---|---|
| Public web pages | Pages a crawler could reach and copy into the index: articles, blogs, reference sites. | Most checkers |
| Publications | Journal articles, books and other published work the tool has access to. | Many checkers. Coverage differs a lot between tools. |
| Previously submitted papers | Work that students handed in through the same platform before. | Some tools, mainly platforms that schools use |
The third row is the one that differs most. As an example, the London School of Economics describes the school platform Turnitin, in its policy on the use of Turnitin, as a service that matches student work against previously submitted student assessments, websites and academic papers. Not every checker has a store of student papers.
The index is the main reason two checkers disagree. If a source is not in the index, the tool cannot match it, however closely your text follows it.
See a plagiarism report for your own text
The process, step by step
1. The text is cleaned up
The checker takes the words out of your file or text box. It usually ignores formatting, and it often ignores capital letters, punctuation and extra spaces. This way a small change in layout does not hide a match.
2. The text is split into short phrases
The checker cuts the text into short runs of words that overlap, a few words at a time. Take the sentence “The industrial revolution changed how people worked and lived.” Cut into runs of five words, it gives “the industrial revolution changed how”, then “industrial revolution changed how people”, and so on to the end.
Short phrases are used because whole sentences rarely match exactly. A copied sentence with one word changed still shares most of its short phrases with the original.
3. Each phrase gets a fingerprint
Comparing text letter by letter against millions of documents would be far too slow. So each phrase is turned into a short code, often called a fingerprint. The same phrase always gives the same code. Looking up a code is very fast.
Many tools keep only a sample of the fingerprints, not every one. That makes the search quicker. It is also one reason a very short copied passage can be missed.
4. The fingerprints are looked up in the index
The sources in the index were split and fingerprinted in the same way when they were added. The checker looks up your fingerprints and gets back a list of sources that share some of them. These are the candidates.
5. Your text is lined up with each candidate
Sharing a few phrases proves little. So the checker compares your text with each candidate more closely. It finds where the shared phrases sit next to each other and joins them into longer passages. Many tools allow small gaps here, so a passage with a few words swapped or removed can still be found.
6. Weak matches are filtered out
Very short matches are usually dropped, because short runs of common words appear everywhere. Some tools let you leave out quoted text or the reference list, or ignore matches below a certain length.
7. The matches are reported
What is left is shown to you: the matching passages highlighted in your text, the source for each, and a score. The next two sections cover the score and the report.
How the score is worked out
The score is usually the share of your text that sits inside a match. If 200 words of a 2,000-word paper are inside matching passages, the score is about 10%.
The score is simple, which is why it needs care:
- It counts words, not wrongdoing. A cited quotation and a copied sentence add to it in the same way.
- It depends on the settings. Leaving out quotations or the reference list changes the number without changing the paper.
- It depends on the tool. A different index gives different matches, so a different score.
So there is no score that means “safe” everywhere. See how much plagiarism is allowed? for why.
What the report shows
A useful report shows more than a number. Look for three things:
- The matching passages, highlighted in your own text.
- The source of each match, with a link, so you can open it and compare.
- The share each source contributes, so you can see whether the matches come from one place or many.
The link to the source matters most. Without it you only know that something matched. With it you can see what matched and decide whether your paper credits it. The Plagiarism Checker Plus report links every matching passage to its source.
For how to go through a report, see how to interpret a plagiarism report. For the full method of checking your own draft, see how to check for plagiarism.
What a plagiarism checker can miss
The method looks for shared wording in stored sources. Anything that is not shared wording, or not in the store, is hard or impossible for it to see.
- Reworded text. When a passage is put fully into different words, few phrases match. Some tools catch close rewording, but none do so reliably. The idea still came from the source and still needs a citation. See is paraphrasing plagiarism?
- Translated text. A passage translated from another language shares no words with its source. The LSE policy linked above lists this as a limit of the tool it uses.
- Sources that are not in the index. A printed book, a paper behind a login, a lecture handout, a page published last week, or a friend’s essay that was never submitted anywhere.
- Ideas and structure. Taking someone’s argument, outline or findings without their words leaves nothing to match.
- Images, charts and equations. A text checker reads text. The same LSE policy notes that its tool cannot identify matches of images, including diagrams and equations.
- Work written by someone else for you. New text that was written to order matches nothing.
These are limits of the tool, not gaps in the rules. Each of these is still plagiarism or misconduct, and readers find them in other ways: by knowing the sources, by noticing a change in voice, by asking the writer about the work. A clean report does not prove a paper is honest. See the types of plagiarism and our plagiarism examples.
What it flags that is not plagiarism
The method also errs the other way. It highlights honest text, because honest text can match too.
- Cited quotations. Quoted words match their source exactly. That is what a quotation is.
- Reference-list entries. Every paper that cites the same source formats it in much the same way.
- Common phrases. “On the other hand” and “plays an important role in” appear in countless texts.
- Names and terms. Titles of books and laws, names of methods and organisations.
- Your own earlier work. If you published a text online, a checker can match your new text against it. Whether reuse is allowed is a separate question. See self-plagiarism.
None of these needs rewording. Check that quotations have quotation marks and a citation, and that references are complete. If you need to build one, the citation generator can help.
Are plagiarism checkers accurate?
It depends on what you ask them to do. A checker is a text-matching tool, and it is good at text matching.
- Reliable: finding passages copied word for word, or nearly so, from a source in the tool’s index.
- Less reliable: finding text that has been reworded, translated, or taken from a source the tool has not indexed.
- Not possible: deciding whether a match is plagiarism. The tool cannot see whether you credited the source, or whether you meant to copy.
No checker finds everything, and every checker flags some honest text. No single figure sums this up, because the result depends on the paper, the sources behind it and the tool’s index. Be wary of any tool that claims to catch all plagiarism.
The practical answer is to use a checker for what it does well. Let it find the matching passages. Then read each one next to its source and decide what it needs.
The first check on Plagiarism Checker Plus is free up to 1,000 words without an account. You can also check inside a document with the Google Docs add-on. AI-written text is a different check: see how AI detectors work.