Guide for writers and SEO teams

Duplicate content and SEO

Duplicate content on your own site is normally a question of which URL Google shows, not a penalty. Copying other sites is a different matter. Here is what Google’s documentation says about both, and what to do.

Updated · Plagiarism Checker Plus editorial team

The short answer

Duplicate content is the same or very similar content at more than one URL, and on your own site it does not normally hurt SEO through a penalty. Google groups the duplicate URLs, picks one as the canonical, and shows that one in search results. Google’s page on canonicalization says: “Some duplicate content on a site is normal and it’s not a violation of Google’s spam policies.”

Three things follow from Google’s documentation:

  • Duplicates on your own site are a canonicalization matter. The question is which URL Google shows, and you can tell Google which one you prefer.
  • Copying other sites is a spam-policy matter. Republishing content from other sites without adding original content or value is listed in Google’s spam policies as abusive scraping.
  • Duplicates still have a cost. Google says the same content at many URLs can be a bad user experience and can make it harder to track how your content performs.
Everything this guide says about Google comes from Google Search Central and Google’s help pages, read on 6 October 2026, and each section links the page it relies on. We read the canonicalization documentation, the spam policies, the helpful content page, the e-commerce URL and pagination pages, and the Search Console help page for the Page indexing report. The pages we read do not say how much any of this affects a ranking, so this guide does not either.

What counts as duplicate content

Duplicate content is content that is the same, or nearly the same, at two or more URLs. It helps to split it into two kinds, because Google’s documentation treats them differently.

Duplicates inside your own site

These are usually caused by how a site is built, not by anyone copying anything. Google’s canonicalization page lists these reasons a site may have duplicate content:

  • Region variants: separate USA and UK URLs with essentially the same content in the same language.
  • Device variants: a page with both a mobile and a desktop version.
  • Protocol variants: the HTTP and HTTPS versions of a site.
  • Site functions: the results of sorting and filtering a category page.
  • Accidental variants: a demo version of the site left open to crawlers.

Other common cases work the same way: URLs with tracking parameters, the www and non-www versions of a domain, printer-friendly pages, and a product page for each size or colour.

Two cases are often called duplicates when they are not. Google’s pagination page says URLs in a paginated sequence are treated as separate pages. And the canonicalization page says different language versions of a page are duplicates only if the primary content is in the same language.

Duplicates across different sites

  • Syndication. Your article is republished by a partner with your permission.
  • Scraped copies. Another site takes your pages without permission.
  • Copied manufacturer or merchant text. Many shops and affiliate sites publish the same product description.
  • Plagiarised articles. A writer copies passages from another site, or rewrites them lightly, and the article is published as new. See what is plagiarism for the definition.

What Google does with duplicates

Google’s canonicalization page describes the process. When Google indexes a page, it works out the primary content. If it finds several pages where the primary content is the same or very similar, it clusters them together. It then chooses the page that is “objectively the most complete and useful for search users” and marks it as the canonical. The same page says:

  • Google uses the canonical page as the main source to evaluate content and quality.
  • A search result usually points to the canonical page.
  • The canonical page is crawled most regularly. Duplicates are crawled less often.
  • You can indicate which URL you prefer, but that is “a hint, not a rule”. Google may choose a different page.

Google’s page on how to specify a canonical URL adds that choosing a canonical helps search engines consolidate the signals they have for the individual URLs, such as links to them, into a single preferred URL. It also says none of the methods are required, and that a site “will likely do just fine without specifying a canonical preference”.

SituationWhat Google’s documentation saysWhat to do
The same page at several URLs (parameters, sorting, filtering)Google clusters the URLs and picks one canonical. Some duplicate content on a site is normal.Add a rel="canonical" link to the URL you prefer and link to that URL inside your site.
HTTP and HTTPS versionsGoogle prefers HTTPS pages over equivalent HTTP pages as canonical, unless there are issues or conflicting signals.Redirect HTTP to HTTPS. List only HTTPS URLs in your sitemap.
A duplicate page you no longer needA redirect is a strong signal that its target should become canonical. Use it when you want to get rid of a duplicate page.Add a permanent redirect to the page you are keeping.
Product variants (size, colour)For products with a URL per variant, include the canonical product URL on all variant pages. If variants use optional query parameters, use the URL without the parameter as the canonical.Pick one product URL and point every variant at it with rel="canonical".
Paginated listsPages in a sequence are treated as separate pages. Do not use the first page as the canonical for the others.Give each page in the sequence its own canonical URL.
Regional versions in the same languageThey are duplicates if the primary content is in the same language. Use both canonicalization and hreflang.Add hreflang annotations so the right regional URL shows.
Your article is syndicated to partner sitesThe canonical link element is not recommended here. The most effective solution is for partners to block indexing of your content.Ask partners to block indexing of their copy.
Another site copied your page without permissionIn rare situations Google may select the external URL. You can contact the site’s host and file a copyright removal request with Google.Follow the steps under “a copycat website” on Google’s troubleshooting page.
Your page republishes other sites’ content with nothing addedThis is listed as abusive scraping in the spam policies. Sites that violate the policies may rank lower or not appear at all.Rewrite the page with your own material, or remove it.

The row on product variants comes from Google’s page on URL structure for e-commerce sites, the row on paginated lists from the pagination page linked above, and the rows on syndication and copycat sites from Fix canonicalization issues.

What about the “duplicate content penalty”?

Google’s current documentation does not use that phrase. The Google page that does is a Search Central Blog post from 12 September 2008, which is still online. It says: “There’s no such thing as a ‘duplicate content penalty.’ At least, not in the way most people mean when they say that.” The same post says there are penalties related to having the same content as another site, and gives scraping and republishing without added value as examples. That is the same line the current documentation draws.

How to find duplicate or copied text

On your own site

Open the Page indexing report in Google Search Console. Google’s help page for the report lists these reasons a page can be shown as not indexed:

  • Duplicate without user-selected canonical. The page is a duplicate and does not name a preferred canonical. Google has chosen the other page. The help page says this “is not an error, but is working as intended”.
  • Duplicate, Google chose different canonical than user. You marked this page as canonical, but Google thinks another URL makes a better one.
  • Alternate page with proper canonical tag. The page correctly points to a canonical page that is indexed. Nothing to do.

The same help page says having a page marked duplicate or alternate “is usually a good thing”, because it means Google found the canonical page and indexed it. Use the URL Inspection tool to see which URL Google selected for any page.

You can also use Google Search itself. The site: operator limits results to one domain, and putting a phrase inside quotes searches for an exact match. Search for site:example.com "one exact sentence from the page" to see which of your URLs carry that sentence. Google notes that site: results are not always a complete list.

On other sites

Search for one or two exact sentences from your article in quotes, without the site: operator. Results from other domains show where the same wording appears.

In a draft, before it is published

This is the check that is easiest to skip. A freelancer’s article, a guest post or a draft built from research notes can contain copied passages that nobody meant to leave in. Searching sentence by sentence is slow, so this is where a plagiarism checker helps.

Plagiarism Checker Plus compares the text you paste with other sources and links each matching source next to your text. Your first check is free without an account, up to 1,000 words, and a free account includes 3,000 words every month. It checks the text you give it. It does not crawl your website or watch the web for copies of your pages. For how to read the result, see how to interpret a plagiarism report.

How to fix it

Google’s page on how to specify a canonical URL lists three methods “in order of how strongly they can influence canonicalization”:

MethodHow strong Google says it isWhen to use it
RedirectA strong signal that the target of the redirect should become canonical.When you are getting rid of the duplicate page. Google says to use permanent redirects.
rel="canonical" linkA strong signal that the specified URL should become canonical.When both URLs need to stay live. Put it in the head of the page, or send it as an HTTP header for files such as PDFs.
Sitemap inclusionA weak signal that helps the URLs in a sitemap become canonical.As a support for the other two. List only the URLs you want as canonicals.

Google says these methods can stack and become more effective when combined. The same page gives best practices. The ones that come up most often:

  • Do not use noindex to pick a canonical within one site. Google says it does not recommend this because it blocks the page from Search completely, and that rel="canonical" is the preferred solution.
  • Do not use robots.txt for canonicalization.
  • Do not use the URL removal tool for it. Google says it hides all versions of a URL from Search.
  • Do not give conflicting signals, such as one URL in the sitemap and a different one in rel="canonical" for the same page.
  • Use absolute URLs in the rel="canonical" link, and include a self-referencing canonical on the canonical page itself.
  • Link to the canonical URL inside your site, not to a duplicate.

When the answer is to rewrite

Canonical tags and redirects solve URL problems. They do not solve a page whose text was taken from somewhere else. If a page copies a manufacturer’s description or another site’s article, the fix is to write your own. If two of your own pages are grouped as duplicates but should both be indexed, Google’s troubleshooting page says the pages need to be sufficiently different, and that pages split out faster when the difference is clear and significant. It also says Google might hold pages in a duplicate cluster for up to two weeks after you fix the content.

For writers and editors

Most of the fixes above belong to whoever runs the site. The wording belongs to you. Google’s page on creating helpful content asks this question about any page: if the content draws on other sources, does it avoid simply copying or rewriting those sources, and instead provide substantial additional value and originality?

  • Write from your own material first. Your own data, examples, tests and experience cannot match another page.
  • Quote when the exact words matter. Put them in quotation marks, name the source and link it.
  • Paraphrase properly and still cite. Swapping synonyms is not paraphrasing. See how to paraphrase without plagiarizing.
  • Write your own product and category copy. Text supplied by a manufacturer is likely to be on other sites as well.
  • Check AI drafts like any other draft. A draft from an AI tool can repeat existing wording. See does ChatGPT plagiarize?
  • Editors: check before you publish or pay. Read each match and decide whether it is a quotation with a source, a common phrase or a copied passage. A percentage alone does not tell you that.

For a fuller checklist, read how to avoid plagiarism.

FAQ

Questions

Usually not through a penalty. Google’s documentation says some duplicate content on a site is normal and is not a violation of its spam policies. Google picks one URL as the canonical and shows that one. The practical costs Google names are a worse user experience, harder tracking, and crawling time spent on duplicate pages.

Keep reading

Check an article for copied text before you publish

Paste up to 1,000 words for a free check, no account needed. The report links each matching source next to your text, so you can attribute it or rewrite it.