Skip to content
Scheidegger Webpublishing Webpublishing, Switzerland
← All articles

Duplicate content: the myth and the real problem

No, Google does not "penalise" accidental duplicates. It picks one version and ignores the others, and that is sometimes worse: the version it picks is not the one you wanted. What canonicalisation actually means.

2 min read

Few subjects carry as much useless fear as “duplicate content”. You still read that a sentence repeated on two pages would earn a sanction. The documented reality is simpler and more interesting: there is no penalty for ordinary duplicates, there is an arbitration, and the real question is who does the arbitrating.

What Google actually does with duplicates

When several addresses serve the same content, Google does not punish: it groups. Within the group it picks a so-called canonical version, the one that will represent the content in the results, and the others are only crawled occasionally from then on. The sanction exists only for deceptive duplication, the kind covered in the series on punished practices: copying other people’s content at scale. The accidental duplicate is a normal phenomenon the engine handles every day.

Where do these accidental duplicates come from? Rarely from copy-and-paste. They are addresses: the version with and without www, HTTP and HTTPS coexisting, the URL with its trailing slash and without it, the tracking parameters that ad platforms hang onto links. Each variant is, to the engine, one more address serving the same page.

The real problem: losing the arbitration

If you say nothing, Google chooses alone, and the documentation is frank on this point: its choice may not be yours. It is the parameter-laden version that starts showing in the results, or the old address instead of the new one. And the signals dilute while the arbitration is pending: external links pointing half to one variant, half to the other, split their weight between two addresses instead of giving it to one.

The tools for keeping the upper hand number three, and the first two already have their articles: the permanent redirect, which remains the strongest signal when a variant can disappear; internal consistency, all your own links written in the same address form; and the canonical tag, one line in the page’s head that declares “this is the reference version”, for the cases where the variants have to keep existing.

The case that concerns everyone here

One worry comes back with every multilingual quote: “our French and German pages say the same thing, is that duplicate content?” No. A translation is not a duplicate, it is different content for different audiences, and the mechanism designed to connect them is the hreflang described in the part on languages. The only real mistake in this area is the one pointed out there: translating the body of the page while leaving titles and descriptions in the original language, because there, yes, you are manufacturing pages the engine can no longer tell apart.

Who writes these notes

This journal is kept by the workshop that designs and maintains the house’s websites. Everything described here, the Search Console, internal links, the business profile, is part of the work delivered with a site: if you would rather someone took care of it, that is precisely the trade.

Design your site Write to us