Home›Features›Duplicate content checker

Crawl · Feature 04

Duplicate content checker

SiteAuditLint compares the main text of every page to find exact duplicates, near-duplicates and thin content, ignoring the boilerplate that every page shares.

04 Crawl5 min Read timeAll editions Edition
Download SiteAuditLint

04 / Crawl

SiteAuditLint compares the main text of every page to find exact duplicates, near-duplicates and thin content, ignoring the boilerplate that every page shares.

Duplicate content splits ranking signals between URLs and makes search engines choose which version to show. Thin content gives them little reason to rank a page at all. Both are hard to spot by browsing a site, and easy to spot with a crawl that measures content similarity.

1. How duplicate detection works

From HTML to similarity groups

1Extract main textMenus, headers, footers and forms are removed.
2Build signaturesEach page gets a MinHash signature of its text.
3CompareSignatures are compared across the whole crawl.
4GroupIdentical pages and 85%+ matches are grouped.

Comparison uses main content only, so shared templates do not create false matches.

2. Exact duplicate content

Exact duplicates usually come from parameter URLs, printer versions, http and https or www and non-www variants, and CMS archive pages that repeat the same text.

WarningDuplicate content (exact)

The main content of this page is identical to other pages, so search engines must choose which to show.

How to fix: Consolidate the pages, make each one unique, or set a canonical to the preferred version.

3. Near-duplicate content

Near-duplicates are the more common problem: location pages with a swapped city name, product variants with the same description, or tag pages listing the same posts. They compete with each other for the same queries.

OpportunityNear-duplicate content

The main content is about 85% or more similar to other pages, which can make them compete in search.

How to fix: Make each page clearly different or consolidate them into one stronger page.

4. Thin content

WarningThin content (under 200 words)

Pages with very little text rarely satisfy search intent and may be treated as low value.

How to fix: Expand the page with useful, original content or consolidate it into a stronger related page.

5. Choosing the right fix

SituationUsual fix
Same page on several URLs301 redirect to one version, or a canonical tag
Variants that must existCanonical to the main version, unique copy where it matters
Pages targeting the same queryMerge into one stronger page and redirect the rest
Thin pages with valueExpand with unique, useful content
Thin pages without valueNoindex or remove them
Related checks

Duplicate titles and meta descriptions are covered in on-page SEO checks. Often the same pages show up in both lists.

Frequently asked questions

How similar do pages need to be to count as near-duplicates?

About 85 percent or more. SiteAuditLint builds a MinHash signature from each page's main text and groups pages whose signatures are that close.

Does shared navigation make every page look duplicated?

No. Menus, headers, footers and forms are removed before comparison, so only the main content is measured.

What counts as thin content?

An indexable page with fewer than 200 words of main text. It is a warning, because short pages can be valid, but they rarely rank for competitive queries.

Want the background first? Browse the SiteAuditLint Academy, or see all features.