Canonical Tags and Duplicate Content: When They Actually Help
Sep 8, 2026 · 7 min read

Duplicate content is not a penalty. Google does not punish anyone for having two identical pages — it just picks one to show and sets the other aside. The problem is that it may not pick the one you wanted. Meanwhile links, clicks and relevance get split across three or four addresses carrying the same text.
That is what the canonical tag is for. Not to be present everywhere, but to tell Google which URL is the main one.
What a canonical tag actually does
In the page head it looks like this:
``html <link rel="canonical" href="https://example.com/category/t-shirts" /> ``
It means: "I know this page exists, but the canonical version is the one at this address." Google treats it as a strong signal, not an instruction. If the canonical contradicts other signals — a different URL in your sitemap, your internal linking, a redirect — Google can pick a different canonical URL. In Search Console you then see it under "Duplicate, Google chose different canonical than user".
That is the whole difference between a canonical and a redirect. A 301 is a command: this address is gone, go elsewhere. A canonical is a recommendation: both addresses work, index this one.
When a canonical tag really solves duplicate content
A canonical is the right answer when the content is genuinely the same or nearly the same and both URLs have to stay reachable.
- URL parameters.
?utm_source=newsletter,?ref=affiliate,?sort=price-asc— from the index's point of view these are all duplicates. A canonical pointing at the clean address merges them. - Variants of the same product. A shirt in five sizes on five URLs with an identical description belongs on one canonical product page.
- Trailing slash and non-slash, HTTP and HTTPS, www and non-www. Redirects are the proper fix, but a canonical keeps the index tidy until the redirects are in place.
- A homepage on two addresses —
/and/index.phpor/home. - Syndicated content. If a partner site republishes your article, a canonical on their page pointing to your URL says who the source is. Cross-domain canonicals work, but they depend on the other side agreeing to add them.
- Print and AMP versions.
And one thing that gets overlooked: the self-referencing canonical, a page pointing at itself. It is not mandatory, but it is the simplest protection against anyone getting your pages indexed with parameters appended.
When a canonical will not help
This is where most of the damage happens, because the canonical gets used as a universal patch.
The content is actually different. If you give two categories a shared canonical just because their titles look similar, one of them drops out of the index — and with it the queries it used to appear for. Similar titles are not fixed by canonicalisation, they are fixed by rewriting the titles. That is a separate job.
Endless filters and faceted navigation. When filter combinations generate tens of thousands of URLs, a canonical merges them in the index but the crawler still has to walk through them. Your crawl budget gets spent on colour combinations. That situation calls for noindex, blocking in robots.txt, or filters implemented without changing the URL.
Pagination. A canonical from ?page=3 to the first page of the category is an old mistake. Google then does not see the products on page three as part of an indexable structure. Every paginated page should have a self-canonical and clear internal linking.
Canonical plus noindex on the same page. Contradictory instructions: first "index the other one", then "do not index at all". Google picks one, and usually not the one you meant.
A canonical pointing at something that is not indexable — a 404, a redirect, a page with noindex, or a URL blocked in robots.txt. The signal gets thrown away.
Content copied onto other people's domains. A canonical on your site has no influence over what someone else does on theirs.
Canonicalisation errors I find most often
| Error | What happens | |---|---| | Two canonical tags on one page | Google ignores both | | Canonical in <body> instead of <head> | It does not count | | Relative URL (/t-shirts) instead of absolute | Risky with multiple domains and variants | | Every subpage canonicalised to the homepage | The whole site drops out of the index except the front page | | Canonical pointing at the HTTP version on an HTTPS site | Split signals | | Sitemap canonical differs from the one in the HTML | Contradictory signals, Google decides on its own | | Canonical injected by JavaScript | Unreliable, depends on rendering | | Translations without a self-canonical alongside hreflang | Language versions merge into each other |
A category of its own is plugin on top of plugin. On WordPress the canonical can be set by an SEO plugin, the theme and an ecommerce extension at the same time. The result is a duplicate tag, or a tag pointing at the old URL structure from before the redesign.
How to check what you have
Three steps you can get through in half an hour:
- Search Console → Page indexing. Look at the counts for "Alternate page with proper canonical tag" (that one is fine) and "Duplicate, Google chose different canonical than user" (that one is a problem). For the second state, write down example URLs.
- URL Inspection. Paste a specific address and, in the indexing section, you get "User-declared canonical" and "Google-selected canonical". When they differ, your canonical is losing to other signals.
- The source code. In the page source, search for
rel="canonical". If it appears twice, or not at all, you have your answer.
When Google's chosen canonical differs from yours, a louder canonical will not help — aligning everything else will. Put only canonical URLs in the sitemap. Point internal links at the canonical version, not at the variant with a parameter. And check whether the canonical page is not in fact the weaker of the pair — Google often picks the one that more internal links point to.
Canonicalisation in an ecommerce store: the practical calls
For online stores it comes down to four rules:
- Sorting and display (
?sort=,?view=) — canonical to the clean category. - Filters that match real search demand (for example "white men's t-shirts") — a separate indexable page with its own copy and its own canonical.
- Filters with no demand, and their combinations —
noindex, or no crawlable links at all. - Pagination — a self-canonical on every page.
The line between the second and the third rule cannot be guessed from a desk. It comes from what people actually type into search, and you can read that out of Search Console data.
How I help with it
When I crawl a site I check the canonical on every public page: whether it exists, whether it sits in the head, whether there is exactly one, whether it is absolute, whether the target URL returns a 200, and whether it conflicts with noindex or with the sitemap. You do not get the finding as advice to "set up canonical URLs" — you get a specific list of addresses and the value that belongs there. Fixes are deployed through the WordPress plugin, the API, or a CSV export, and you only approve them.
One thing matters with canonicalisation: the effect does not show up in days. Google has to recrawl the pages and reconsider which version it treats as canonical. On larger sites that means weeks. So I measure the 28 days before deployment and the 28 days after — otherwise there is no way to separate a real change from a seasonal swing.
Checklist
- Every indexable page has one canonical, in
<head>, with an absolute URL. - The canonical points at an address that returns 200 and has no
noindex. - The sitemap contains canonical URLs only.
- Paginated pages have a self-canonical, not a link to page one.
- Filter combinations are set to noindex instead of merely being canonicalised.
- Internal links point at the canonical version of the address.
- In Search Console, the number of URLs in "Duplicate, Google chose different canonical" is going down.
The canonical tag is a precise tool for one job: merging signals from addresses with the same content. Everywhere else it is just a way to drop your own pages out of the index.
#seoobsah
This is written by a tool you can buy
The article was proposed and written by Seonal — the same one that finds the errors on your site, fixes them and measures the result. The audit is free.