Duplicate content in ecommerce means the same product or category content is reachable at more than one URL, or the same text appears on your store and on other retailers’ sites. It isn’t a penalty. Google says some duplication is normal and not a spam violation, but it still picks one URL to show, splits signals like links across the copies, and burns crawl time on pages you never meant to create.
So don’t panic. Find where each copy comes from, then give Google one clear answer. Below are the sources I check in every store audit, the fix for each, and a decision table you can work through line by line.
Is There Really a Duplicate Content Penalty?
No. Google’s canonicalization documentation states that “some duplicate content on a site is normal and it’s not a violation of Google’s spam policies.” Back in 2008 Google even published a post titled Demystifying the “duplicate content penalty” to bury the myth.
There’s one real exception. Google’s spam policies cover scraped content: lifting other sites’ pages, tweaking a few words or none, and republishing them to manipulate rankings. That’s deception, and it can earn a manual action. A size filter spawning a second boots URL is a different universe.
What actually happens with ordinary store duplication is quieter. Google groups the copies, chooses a canonical and shows that one. Leave it guessing and, in my experience, it sometimes crowns the URL with a tracking parameter or the long collection path, which is exactly the outcome worth preventing.
Four Questions I Ask Before Picking a Fix
Every fix below grows out of four questions. I answer them for each pattern before touching a template or robots.txt.
- Do shoppers need both URLs? A colour variant helps someone buy. An
http://copy of your homepage helps nobody. - Is the content identical, or only similar? Red and blue shirts share most text but show different photos and stock.
- Is the copy on your domain or someone else’s? Internal copies are a technical job. Cross-site copies are a writing job.
- How many URLs does the pattern create? Ten can live with a canonical tag. Ten thousand filter URLs need the crawl stopped.
Short version? Redirect what shoppers don’t need. Canonicalize what they do need but search doesn’t. Block patterns that explode.
What Causes Duplicate Content on Ecommerce Sites?
Stores breed duplicates faster than any other kind of site I work on, because the platform builds a fresh URL for every filter, sort order and navigation path a shopper might take. These are the eight sources I check, roughly ranked by how often they bite.
Faceted Filter URLs
Colour, size, brand and price filters each add a parameter, and they stack. One category can produce thousands of slightly different slices of the same products. Google’s faceted navigation guidance explains that crawlers can’t judge those URLs without fetching them, so they tend to fetch enormous numbers.
This is the biggest source by volume. It’s mostly a crawl-waste issue rather than a ranking one, and I’ve covered the “should this filter become a real page” call in my category page SEO guide, so I’ll leave it there.
Sort Order and Pagination
A ?sort=price-asc URL lists identical products in a different order. True duplicate. Canonicalize it to the clean category.
Pagination is where people over-fix. Page 2 shows different products than page 1, so it isn’t a copy at all. Google’s pagination advice is to give every page a self-referencing canonical rather than pointing the whole series at page 1.
Product Variants With Their Own URLs
Google’s ecommerce URL structure guide accepts both /t-shirt/green and /t-shirt?color=green. When variants use an optional parameter, Google says to make the parameter-free URL canonical.
My rule is about demand. A colourway people search for by name can keep its own canonical and unique copy. A size can’t. Point it at the parent.
Collection-Scoped Product URLs
Plenty of platforms let one product live at /products/boot and also at /collections/winter/products/boot. Put that boot in five collections and you’ve got six addresses. The platform usually adds a canonical to the short path, yet the theme may still link to the long one from every grid.
I prefer fixing it upstream. When every internal link points at the canonical product URL, Google rarely meets the duplicates. Shopify has quirks of its own, which I tested in my Shopify collection page SEO post.
Manufacturer Descriptions
This one crosses domains. Fifty retailers paste the same supplier blurb, Google sees fifty near-identical pages and usually shows a handful. No penalty. Just no reason to pick you.
Honestly, no tag solves it. Original copy does. Start with best sellers and high-margin items, then follow my guide on how to write product descriptions for SEO.
Printer Pages, Session IDs and Tracking Codes
Older platforms still generate printer-friendly views and stuff session IDs into URLs. Marketing layers ?utm_source= and affiliate codes on top. Google’s URL guide warns against internally linking to session IDs, tracking codes and similar temporary parameters.
Keep them off internal links. Store session state in cookies. Make sure every leftover variant carries a canonical to the clean address.
Protocol, Host and Case Variants
Think http versus https, www versus bare domain, trailing slash or none, capitals or lowercase. Each can count as a separate URL. Pick one version and 301 everything else to it in a single hop; my URL structure SEO guide lists the variants Google treats as distinct.
International Copies
A store selling in both the US and UK often runs two English sites with near-identical text. Google’s canonicalization docs name this case directly, and add that different language versions only count as duplicates when the main content shares a language.
For same-language regional stores, use hreflang so each country gets its own URL. Never canonicalize the UK store to the US one. Do that and Google can drop the UK pages from UK results.
Which Fix Should You Use for Each Case?
Google’s guide to consolidating duplicate URLs ranks the tools. Redirects and rel="canonical" are both strong signals. Sitemap inclusion is weak. Here’s how I choose.
If shoppers never need the duplicate, then 301 it. Protocol, host and retired paths belong here, and my 301 vs 302 redirects guide explains why the permanent status matters.
If shoppers need the URL but search doesn’t, keep it live and canonicalize it. Sorts, single filters, tracking parameters and minor variants sit in this branch.
If the pattern creates thousands of URLs, stop the crawl with robots.txt and treat the canonical as backup. Google’s faceted navigation page calls canonicals generally less effective over time for this case.
If the copy lives on another site, nothing in your <head> helps much. Rewrite.
The Noindex Caveat
People grab noindex because it feels decisive. Google’s consolidation doc advises against using it to choose a canonical within one site, since it drops the page from Search entirely along with its signals.
I reserve noindex for pages that should never appear anywhere in results, such as internal search or account screens. Stacking noindex and a canonical on one URL hands Google two conflicting instructions. Send one.
Robots.txt Controls Crawling, Not Indexing
This misunderstanding does more damage than anything else in my audits. Robots.txt tells Googlebot not to fetch a URL; it doesn’t remove an indexed page, and Google can still index a blocked address if other pages link to it.
Worse, Google can’t read a canonical or noindex on a page it may not fetch. The consolidation doc says outright: don’t use robots.txt for canonicalization. Already have indexed filter URLs? Let Google crawl them, see the noindex or canonical, wait until they drop, and only then add the block.
What Replaced the URL Parameters Tool?
Nothing, officially. In March 2022 Google announced, in a Search Central post titled “Spring cleaning: the URL Parameters tool,” that the Search Console setting was going away. Google said only about 1% of the parameter configurations site owners had set were useful for crawling.
Parameter handling now lives on your site. I pull three levers together:
- Clean internal links. Navigation and product grids never link to tracking, session or sort parameters.
- Canonicals on parameter URLs whose content matches a clean URL.
- Robots.txt rules for patterns that should never be crawled, like sort orders and multi-filter combinations.
Then check the page indexing report in Search Console. Handled duplicates appear as Alternate page with proper canonical tag. That’s healthy. Leave it.
Duplicate Content Ecommerce Decision Table
Here’s the whole post in one grid. Find the source, check the condition, apply the fix.
| Duplicate source | Condition | Fix |
|---|---|---|
| Filter URLs | Low value, thousands of combinations | Robots.txt disallow; canonical as backup |
| Filter URLs | Real search demand | Promote to its own category page |
| Sort parameters | Same products, reordered | Canonical to clean category URL |
| Pagination | Page 2+ lists different products | Self-referencing canonical per page |
| Variants | Size or minor option | Canonical to parent product |
| Variants | Colourway with its own demand | Self canonical plus unique copy |
| Collection-scoped product URL | Same product, longer path | Link to canonical URL; keep canonical tag |
| Tracking or session IDs | Any | Remove from internal links; canonical to clean URL |
| http, www, slash, case variants | Any | Single-hop 301 to one version |
| Regional English copies | Different country, same language | Hreflang, each page self-canonical |
| Manufacturer description | Same text on many retailers | Rewrite priority products |
| Discontinued product | No longer sold | See the out-of-stock guide below |
That last row deserves its own flowchart. My out of stock products SEO guide covers when to keep, redirect or remove those pages.
Edge Cases That Change the Answer
A few situations break the neat rules, and each one shows up in store audits more often than you would expect.
- The canonical target redirects. Mixed signals. Point every canonical straight at the final URL returning 200.
- The canonical product is sold out but a variant isn’t. Shoppers land on a dead end, so rethink which URL leads.
- JavaScript injects the canonical. Google can read it after rendering, but I’d still ship it in raw HTML.
- A replatform changed every product path. That’s a redirect-mapping project, not a canonical tweak.
When Should You Get Help?
With a few hundred products and one or two patterns, the table above is enough to do this yourself. It gets harder when filter URLs already run into the tens of thousands, when a migration left old paths live, or when several regional stores share one catalog.
That’s the point where I’d bring in a technical specialist, because one careless robots.txt line can hide an entire category. Our ecommerce SEO service begins with a crawl of exactly these patterns, and a free audit will show which ones your store has. For how pages should be arranged in the first place, read my ecommerce site structure guide.
Frequently Asked Questions
Does Duplicate Content Hurt Ecommerce SEO?
It doesn’t trigger a penalty, but it can still cost you. Google shows one URL from each duplicate group, so your preferred version may lose out. Links get split across copies, and crawl time drifts toward filter and parameter URLs instead of new products.
Should I Use Canonical Tags or 301 Redirects for Duplicate Products?
Use a 301 when shoppers never need the duplicate, like an http or non-www copy. Use a canonical when the URL has a shopping job, like a sort order or size variant, but shouldn’t rank alone. Google treats both as strong signals.
Is Using Manufacturer Product Descriptions Duplicate Content?
Yes, it’s cross-site duplication, though not spam by itself. The problem is competitive: your page looks like every other retailer’s, so Google has little reason to prefer it. Rewrite descriptions for your top sellers and highest-margin products first.
Can I Block Duplicate URLs With Robots.txt?
You can block crawling, not indexing. Google won’t see a canonical or noindex on a blocked URL, and it can still index a blocked page that others link to. Use robots.txt for crawl-heavy patterns like filters, never to choose canonicals.
Last updated: September 2026 by Mizanur Rahman



