04 / Crawl
SiteAuditLint compares the main text of every page to find exact duplicates, near-duplicates and thin content, ignoring the boilerplate that every page shares.
Duplicate content splits ranking signals between URLs and makes search engines choose which version to show. Thin content gives them little reason to rank a page at all. Both are hard to spot by browsing a site, and easy to spot with a crawl that measures content similarity.
1. How duplicate detection works
From HTML to similarity groups
Comparison uses main content only, so shared templates do not create false matches.
2. Exact duplicate content
Exact duplicates usually come from parameter URLs, printer versions, http and https or www and non-www variants, and CMS archive pages that repeat the same text.
WarningDuplicate content (exact)
The main content of this page is identical to other pages, so search engines must choose which to show.
How to fix: Consolidate the pages, make each one unique, or set a canonical to the preferred version.
3. Near-duplicate content
Near-duplicates are the more common problem: location pages with a swapped city name, product variants with the same description, or tag pages listing the same posts. They compete with each other for the same queries.
OpportunityNear-duplicate content
The main content is about 85% or more similar to other pages, which can make them compete in search.
How to fix: Make each page clearly different or consolidate them into one stronger page.
4. Thin content
WarningThin content (under 200 words)
Pages with very little text rarely satisfy search intent and may be treated as low value.
How to fix: Expand the page with useful, original content or consolidate it into a stronger related page.
5. Choosing the right fix
| Situation | Usual fix |
|---|---|
| Same page on several URLs | 301 redirect to one version, or a canonical tag |
| Variants that must exist | Canonical to the main version, unique copy where it matters |
| Pages targeting the same query | Merge into one stronger page and redirect the rest |
| Thin pages with value | Expand with unique, useful content |
| Thin pages without value | Noindex or remove them |
Duplicate titles and meta descriptions are covered in on-page SEO checks. Often the same pages show up in both lists.
Frequently asked questions
How similar do pages need to be to count as near-duplicates?
About 85 percent or more. SiteAuditLint builds a MinHash signature from each page's main text and groups pages whose signatures are that close.
Does shared navigation make every page look duplicated?
No. Menus, headers, footers and forms are removed before comparison, so only the main content is measured.
What counts as thin content?
An indexable page with fewer than 200 words of main text. It is a warning, because short pages can be valid, but they rarely rank for competitive queries.
Want the background first? Browse the SiteAuditLint Academy, or see all features.