Home›Checklists›Indexability Checklist

Free SEO checklist · 13 sections · 13 checks

Indexability Checklist

An indexability checklist walks through every signal that decides whether a page can appear in search results: status code, meta robots, X-Robots-Tag, canonical, robots.txt, internal links, sitemap inclusion, rendering, and duplication.

13 Checks7 High priority13 SectionsPDF · Word · Excel Formats
Download checklist

SiteAuditLint checklist library

What this checklist covers

An indexability checklist walks through every signal that decides whether a page can appear in search results: status code, meta robots, X-Robots-Tag, canonical, robots.txt, internal links, sitemap inclusion, rendering, and duplication.

It is organized into 13 sections: HTTP Status, Noindex, X-Robots-Tag, Canonical, Robots.txt, Internal Links, XML Sitemap, Redirects, JavaScript, Duplicates, Soft 404s, Orphan Pages and Search Console. Each check lists exactly what to verify and a priority, so you can work through the highest impact items first.

For background on the concepts behind these checks, see Indexing explained, Crawlability vs indexability, Noindex issues.

Who it is for:

SEOsdevelopers

Download the indexability checklist

Use the interactive version below, or download it to share with your team, attach to tickets, or work through offline.

Indexability Checklist

Work through each check

0 of 13 checks complete

HTTP Status

Noindex

X-Robots-Tag

Canonical

Robots.txt

XML Sitemap

Redirects

JavaScript

Duplicates

Soft 404s

Orphan Pages

Search Console

Progress is saved in this browser only.

Detailed explanations

Each section below explains what to check, why it matters, the problems you will usually find, how to fix them, and how to confirm the fix.

HTTP Status

What to check:

  • Check status. Page returns 200.

Why it matters: HTTP status codes tell crawlers whether a URL is live, moved, missing, or broken. A 200 response is required for indexing. Unexpected 4xx and 5xx responses remove pages from search results and waste crawl budget, while soft 404s confuse search engines about which pages hold real content.

Common problems:

  • Internal links pointing to 404 pages after content is deleted
  • Soft 404s, where empty or error pages return 200
  • Intermittent 5xx errors on heavy templates or under load
  • Single page apps returning 200 for routes that do not exist

How to fix it: Restore or redirect removed URLs that still receive links, return a true 404 or 410 for content that is gone, and investigate server logs for the cause of any 5xx responses.

How to verify: Crawl the site and filter by status code. Every URL in the sitemap and main navigation should return 200. Recheck flagged URLs individually with a header checker.

Learn more: Status codes lesson, HTTP status codes and SEO, HTTP 4xx issue, HTTP 5xx issue.

How SiteAuditLint helps: it checks http status across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Noindex

What to check:

  • Check meta robots. No noindex in meta robots.

Why it matters: A page can be crawled but still excluded from search results. Indexability depends on several signals working together: status code, meta robots, the X-Robots-Tag header, canonical tags, and robots.txt. A single conflicting signal is enough to remove a page from Google.

Common problems:

  • A noindex left in place after a staging build goes live
  • X-Robots-Tag noindex sent by a server or CDN rule nobody remembers
  • Canonical tags pointing to a different URL than intended
  • Pages blocked in robots.txt, so Google never sees the noindex or canonical

How to fix it: Remove noindex directives from pages that should rank, align canonical tags with the URL you want indexed, and make sure robots.txt does not block indexable content.

How to verify: Check the page source and response headers for every key template, then use Search Console URL Inspection to confirm Google sees the page as indexable.

Learn more: Indexing explained, Indexing tests, Noindex issues and fixes, Noindex issue.

How SiteAuditLint helps: it checks noindex across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

X-Robots-Tag

What to check:

  • Check HTTP headers. No noindex in X-Robots-Tag.

Why it matters: A page can be crawled but still excluded from search results. See the earlier section on this topic for common problems.

How to verify: Check the page source and response headers for every key template, then use Search Console URL Inspection to confirm Google sees the page as indexable.

How SiteAuditLint helps: it checks x-robots-tag across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Canonical

What to check:

  • Check canonical. Canonical is self-referencing or points to the preferred URL.

Why it matters: The canonical tag tells search engines which version of a page is the preferred one when duplicates or variants exist. It consolidates ranking signals onto one URL. A broken canonical can quietly deindex the page you care about most.

Common problems:

  • Canonicals pointing to staging or development domains
  • Relative canonical URLs that resolve incorrectly
  • Canonicals pointing to redirected, 404, or noindexed URLs
  • Every paginated or variant page canonicalized to page one

How to fix it: Use absolute canonical URLs on the production domain, make indexable pages self canonical, and point duplicates at a live, indexable 200 URL.

How to verify: Crawl the site and compare the canonical URL with the page URL. Flag any canonical that is non self referencing, redirects, returns an error, or is noindexed.

Learn more: Canonicals lesson, Canonical tag issues and fixes, Missing canonical issue, Canonicalised URL issue.

How SiteAuditLint helps: it checks canonical across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Robots.txt

What to check:

  • Check robots.txt. URL is not disallowed.

Why it matters: Robots.txt controls which paths crawlers may request. It is the fastest way to block an entire site by accident. It also governs access for AI crawlers, so it is now part of your AI search visibility strategy as well as your traditional SEO.

Common problems:

  • A staging Disallow: / rule deployed to production
  • CSS and JavaScript blocked, stopping Google from rendering pages
  • Overly broad wildcard patterns blocking important directories
  • Missing or outdated sitemap declaration

How to fix it: Keep robots.txt minimal, block only what you truly need to keep out of crawlers, allow rendering resources, and declare your XML sitemap with an absolute URL.

How to verify: Fetch /robots.txt on production, test a list of key URLs against the live rules for Googlebot and each AI user agent, and confirm the file returns 200.

Learn more: Robots.txt lesson, Robots.txt and AI crawlers guide, Blocked by robots.txt issue, Missing robots.txt issue.

How SiteAuditLint helps: it checks robots.txt across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

What to check:

  • Check internal links. Page is linked from crawlable pages.

Why it matters: Internal links distribute authority, define site architecture, and help crawlers discover content. Broken links waste crawl budget and frustrate users, while orphan pages are almost invisible to search engines.

Common problems:

  • Links to deleted pages returning 404
  • Links pointing to redirected URLs instead of final destinations
  • Orphan pages with no internal links
  • Generic anchor text like click here on important links

How to fix it: Fix or remove broken links, update links to point at final URLs, link to orphan pages from relevant hubs, and use descriptive anchor text.

How to verify: Crawl the site and review the broken links, redirecting links, and inlinks reports. Compare sitemap URLs against crawled URLs to find orphans.

Learn more: Internal links lesson, Fix broken internal links, Orphan pages, Internal link analysis.

How SiteAuditLint helps: it checks internal links across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

XML Sitemap

What to check:

  • Check sitemap inclusion. Page is listed in the sitemap.

Why it matters: An XML sitemap is a list of URLs you want search engines to find and index. It helps discovery of new and deep content and acts as a canonical signal. A sitemap full of redirects, errors, or noindexed URLs weakens trust in that signal.

Common problems:

  • Sitemaps containing redirected, 404, or noindexed URLs
  • Non canonical URL variants listed instead of the canonical version
  • Sitemap not referenced in robots.txt or submitted in Search Console
  • Stale sitemaps that are not regenerated on publish

How to fix it: Generate the sitemap automatically from indexable canonical URLs only, keep lastmod accurate, reference it in robots.txt, and submit it in Search Console.

How to verify: Crawl the sitemap URL list on its own. Every entry should return 200, be indexable, and be self canonical.

Learn more: XML sitemaps lesson, XML sitemap errors and fixes, Sitemap non-200 issue, Noindex URL in sitemap issue.

How SiteAuditLint helps: it checks xml sitemap across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Redirects

What to check:

  • Check redirects. URL does not redirect.

Why it matters: Redirects pass users and ranking signals from an old URL to a new one. Each extra hop adds latency and increases the chance that signals are lost or crawlers give up. Temporary redirects used for permanent moves can keep the old URL indexed.

Common problems:

  • Redirect chains of three or more hops left behind by repeated migrations
  • Redirect loops that make a URL unreachable
  • 302 redirects used for permanent URL changes
  • Internal links still pointing at redirecting URLs

How to fix it: Point every redirect directly at its final destination with a 301 or 308, break any loops, and update internal links so they reference the final URL rather than relying on the redirect.

How to verify: Crawl with redirect following enabled and export the redirect chains report. Each redirecting URL should resolve in one hop to a URL returning 200.

Learn more: Redirects lesson, Testing redirects, How to find and fix redirect problems, Redirect chain issue.

How SiteAuditLint helps: it checks redirects across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

JavaScript

What to check:

  • Check rendering. Main content renders for Googlebot.

Why it matters: Google renders JavaScript, but rendering is delayed, resource limited, and not guaranteed. Many AI crawlers do not render JavaScript at all. Content, links, and SEO tags that only appear after scripts run are at risk of being missed.

Common problems:

  • Main content injected only on the client
  • Links implemented as buttons or onClick handlers
  • Titles, canonicals, or meta robots changed by JavaScript
  • Single page app routes returning 200 for missing pages

How to fix it: Server render or statically generate critical content and tags, use real anchor links with URLs, and return proper status codes from the server.

How to verify: Compare raw HTML with rendered HTML for each template. Titles, H1s, main content, canonicals, and links should be present in the raw response.

Learn more: JavaScript SEO course, Rendering lesson, JavaScript links lesson, JavaScript SEO guide.

How SiteAuditLint helps: it checks javascript across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Duplicates

What to check:

  • Check duplicates. Page is not a near duplicate of another URL.

Why it matters: Duplicate and near duplicate URLs split ranking signals and make search engines choose a version for you. They often come from technical variants rather than copied content.

Common problems:

  • Trailing slash and non slash versions both returning 200
  • Uppercase and lowercase URL variants
  • Tracking and sort parameters creating indexable copies
  • Printer friendly or session ID versions of pages

How to fix it: Pick one URL format, redirect other variants to it, and canonicalize parameter versions that must remain accessible.

How to verify: Crawl and group pages by identical or near identical content hashes and titles, then confirm each group resolves to one indexable URL.

Learn more: Duplicate content SEO, Duplicate content checker, Near duplicate issue, URL parameters issue.

How SiteAuditLint helps: it checks duplicates across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Soft 404s

What to check:

  • Check soft 404. Page has substantive content.

Why it matters: HTTP status codes tell crawlers whether a URL is live, moved, missing, or broken. See the earlier section on this topic for common problems.

How to verify: Crawl the site and filter by status code. Every URL in the sitemap and main navigation should return 200. Recheck flagged URLs individually with a header checker.

How SiteAuditLint helps: it checks soft 404s across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Orphan Pages

What to check:

  • Check orphans. Page has inbound internal links.

Why it matters: Internal links distribute authority, define site architecture, and help crawlers discover content. See the earlier section on this topic for common problems.

How to verify: Crawl the site and review the broken links, redirecting links, and inlinks reports. Compare sitemap URLs against crawled URLs to find orphans.

How SiteAuditLint helps: it checks orphan pages across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Search Console

What to check:

  • Inspect URL. URL Inspection shows indexed or indexable.

Why it matters: Without tracking you cannot measure the impact of launches, migrations, or fixes. Missing tags after a deployment create gaps that cannot be backfilled.

Common problems:

  • Analytics missing on new templates
  • Duplicate tags inflating sessions
  • Conversion events lost after a redesign
  • Search Console not verified on the new domain

How to fix it: Deploy analytics through a tag manager across all templates, verify Search Console, and test conversion events before launch.

How to verify: Use tag assistant and real time reports to confirm tags fire on each template and that conversions record.

Learn more: Tracking and analytics audit, Tracking pixel audit, Tracking missing on page issue, SEO KPIs lesson.

How SiteAuditLint helps: it checks search console across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Common mistakes

Blocking a page in robots.txt and expecting noindex to work

Assuming sitemap inclusion guarantees indexing

Ignoring canonical conflicts

Not inspecting the URL in Search Console

When to run the checklist

  • When a page is not appearing in Google
  • When Search Console reports excluded pages
  • After launches and migrations
  • Before submitting important new pages

How to verify fixes

  1. Save a baselineCrawl the site before making changes so you have a record of every status code, canonical, directive and title.
  2. Fix by priorityStart with High priority checks and issues that affect templates, since one fix there resolves many URLs. See how to prioritize audit findings.
  3. Re-crawl the same scopeUse the same start URL, crawl limit and settings so the results are comparable.
  4. Compare the crawlsConfirm the issue count dropped and that no new problems appeared elsewhere. Audit comparison does this field by field.
  5. Confirm in Search ConsoleUse URL Inspection and the indexing reports to check that Google sees the change. Our indexing tests lesson covers the process.

SiteAuditLint workflow

From manual checklist to automated checks

01CrawlCrawl the whole site, not a sample
02FindSee every affected URL per check
03FixPrioritize by severity and reach
04Re-crawlRun the same scope again
05CompareDiff results against the baseline
06MonitorCatch regressions after each release

Most checks in this list run automatically in a SiteAuditLint crawl. Use audit comparison to diff crawls and scheduled audits to monitor for regressions.