Home›Checklists›Robots.txt Checklist

Free SEO checklist · 10 sections · 12 checks

Robots.txt Checklist

A robots.txt checklist verifies that your robots.txt file is accessible, valid, and allows search engines and the AI crawlers you choose to reach important content, while blocking only what you intend.

12 Checks7 High priority10 SectionsPDF · Word · Excel Formats
Download checklist

SiteAuditLint checklist library

What this checklist covers

A robots.txt checklist verifies that your robots.txt file is accessible, valid, and allows search engines and the AI crawlers you choose to reach important content, while blocking only what you intend.

It is organized into 10 sections: File Accessibility, Syntax, User-agent Rules, Allow and Disallow, Search Crawlers, AI Crawlers, Sitemap Declaration, Wildcards, Testing and Deployment. Each check lists exactly what to verify and a priority, so you can work through the highest impact items first.

For background on the concepts behind these checks, see Robots.txt lesson, Check crawlability for Google, Bing and AI bots, Robots.txt and AI crawlers guide.

Who it is for:

Developerstechnical SEOs

Download the robots.txt checklist

Use the interactive version below, or download it to share with your team, attach to tickets, or work through offline.

Robots.txt Checklist

Work through each check

0 of 12 checks complete

File Accessibility

Syntax

User-agent Rules

Allow and Disallow

Search Crawlers

AI Crawlers

Sitemap Declaration

Wildcards

Testing

Deployment

Progress is saved in this browser only.

Detailed explanations

Each section below explains what to check, why it matters, the problems you will usually find, how to fix them, and how to confirm the fix.

File Accessibility

What to check:

  • Check file location. File is at the root and returns 200.
  • Check file size. File is under 500 KiB.

Why it matters: Robots.txt controls which paths crawlers may request. It is the fastest way to block an entire site by accident. It also governs access for AI crawlers, so it is now part of your AI search visibility strategy as well as your traditional SEO.

Common problems:

  • A staging Disallow: / rule deployed to production
  • CSS and JavaScript blocked, stopping Google from rendering pages
  • Overly broad wildcard patterns blocking important directories
  • Missing or outdated sitemap declaration

How to fix it: Keep robots.txt minimal, block only what you truly need to keep out of crawlers, allow rendering resources, and declare your XML sitemap with an absolute URL.

How to verify: Fetch /robots.txt on production, test a list of key URLs against the live rules for Googlebot and each AI user agent, and confirm the file returns 200.

Learn more: Robots.txt lesson, Robots.txt and AI crawlers guide, Blocked by robots.txt issue, Missing robots.txt issue.

How SiteAuditLint helps: it checks file accessibility across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Syntax

What to check:

  • Validate syntax. Directives parse correctly with no typos.

Why it matters: Robots.txt controls which paths crawlers may request. See the earlier section on this topic for common problems.

How to verify: Fetch /robots.txt on production, test a list of key URLs against the live rules for Googlebot and each AI user agent, and confirm the file returns 200.

How SiteAuditLint helps: it checks syntax across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

User-agent Rules

What to check:

  • Check user-agent groups. Groups are correctly scoped.

Why it matters: Robots.txt controls which paths crawlers may request. See the earlier section on this topic for common problems.

How to verify: Fetch /robots.txt on production, test a list of key URLs against the live rules for Googlebot and each AI user agent, and confirm the file returns 200.

How SiteAuditLint helps: it checks user-agent rules across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Allow and Disallow

What to check:

  • Check disallow rules. No important paths blocked.
  • Check CSS and JS access. Rendering resources are not blocked.

Why it matters: Robots.txt controls which paths crawlers may request. See the earlier section on this topic for common problems.

How to verify: Fetch /robots.txt on production, test a list of key URLs against the live rules for Googlebot and each AI user agent, and confirm the file returns 200.

How SiteAuditLint helps: it checks allow and disallow across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Search Crawlers

What to check:

  • Check Googlebot and Bingbot. Search crawlers can access key sections.

Why it matters: Search engines can only rank what they can discover. If crawlers cannot follow links to a URL, that page will rarely be indexed, no matter how good its content is. Crawl depth also signals importance: pages buried many clicks deep get crawled less often.

Common problems:

  • Navigation built with JavaScript click handlers instead of anchor links
  • Important pages only reachable through internal search or forms
  • Pagination or filters creating near infinite URL spaces that waste crawl budget
  • Key pages sitting five or more clicks from the homepage

How to fix it: Use standard <a href> links for all navigation, add contextual links from high authority pages to deep content, and constrain parameter and filter URLs so crawlers spend time on pages that matter.

How to verify: Run a full crawl, compare discovered URLs with your sitemap and analytics landing pages, and review the crawl depth report for any important URL deeper than three clicks.

Learn more: Crawling explained, What is a website crawler, Crawlability vs indexability, Deep page issue.

How SiteAuditLint helps: it checks search crawlers across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

AI Crawlers

What to check:

  • Check AI crawler rules. GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot rules match policy.

Why it matters: Search engines can only rank what they can discover. See the earlier section on this topic for common problems.

How to verify: Run a full crawl, compare discovered URLs with your sitemap and analytics landing pages, and review the crawl depth report for any important URL deeper than three clicks.

How SiteAuditLint helps: it checks ai crawlers across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Sitemap Declaration

What to check:

  • Declare sitemap. Sitemap directive uses an absolute URL.

Why it matters: An XML sitemap is a list of URLs you want search engines to find and index. It helps discovery of new and deep content and acts as a canonical signal. A sitemap full of redirects, errors, or noindexed URLs weakens trust in that signal.

Common problems:

  • Sitemaps containing redirected, 404, or noindexed URLs
  • Non canonical URL variants listed instead of the canonical version
  • Sitemap not referenced in robots.txt or submitted in Search Console
  • Stale sitemaps that are not regenerated on publish

How to fix it: Generate the sitemap automatically from indexable canonical URLs only, keep lastmod accurate, reference it in robots.txt, and submit it in Search Console.

How to verify: Crawl the sitemap URL list on its own. Every entry should return 200, be indexable, and be self canonical.

Learn more: XML sitemaps lesson, XML sitemap errors and fixes, Sitemap non-200 issue, Noindex URL in sitemap issue.

How SiteAuditLint helps: it checks sitemap declaration across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Wildcards

What to check:

  • Test wildcards. * and $ patterns match only intended URLs.

Why it matters: Robots.txt controls which paths crawlers may request. See the earlier section on this topic for common problems.

How to verify: Fetch /robots.txt on production, test a list of key URLs against the live rules for Googlebot and each AI user agent, and confirm the file returns 200.

How SiteAuditLint helps: it checks wildcards across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Testing

What to check:

  • Test key URLs. Key URLs tested against live rules.

Why it matters: Robots.txt controls which paths crawlers may request. See the earlier section on this topic for common problems.

How to verify: Fetch /robots.txt on production, test a list of key URLs against the live rules for Googlebot and each AI user agent, and confirm the file returns 200.

How SiteAuditLint helps: it checks testing across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Deployment

What to check:

  • Check after deploy. Production file is not the staging file.

Why it matters: Launches often ship staging settings to production. A single forgotten noindex or Disallow can keep a new site out of search for weeks.

Common problems:

  • Staging noindex left on production
  • Robots.txt still disallowing everything
  • Canonicals pointing to the staging domain
  • Analytics not installed

How to fix it: Use a launch runbook that removes staging protections, swaps domains in canonicals and sitemaps, and verifies tracking.

How to verify: Crawl production immediately after launch and compare against the final staging crawl.

Learn more: Pre-launch website checklist, Run your first audit, Noindex issues.

How SiteAuditLint helps: it checks deployment across every crawled URL instead of one page at a time and lists exactly which URLs are affected.

Common mistakes

Shipping a staging Disallow: / to production

Blocking CSS and JavaScript

Using robots.txt to remove pages from the index

Copying AI crawler blocks without a policy decision

Forgetting that rules are case sensitive

When to run the checklist

  • Before and after every deployment
  • Before launching a site
  • When setting an AI crawler policy
  • When pages unexpectedly drop from search

How to verify fixes

  1. Save a baselineCrawl the site before making changes so you have a record of every status code, canonical, directive and title.
  2. Fix by priorityStart with High priority checks and issues that affect templates, since one fix there resolves many URLs. See how to prioritize audit findings.
  3. Re-crawl the same scopeUse the same start URL, crawl limit and settings so the results are comparable.
  4. Compare the crawlsConfirm the issue count dropped and that no new problems appeared elsewhere. Audit comparison does this field by field.
  5. Confirm in Search ConsoleUse URL Inspection and the indexing reports to check that Google sees the change. Our indexing tests lesson covers the process.

SiteAuditLint workflow

From manual checklist to automated checks

01CrawlCrawl the whole site, not a sample
02FindSee every affected URL per check
03FixPrioritize by severity and reach
04Re-crawlRun the same scope again
05CompareDiff results against the baseline
06MonitorCatch regressions after each release

Most checks in this list run automatically in a SiteAuditLint crawl. Use audit comparison to diff crawls and scheduled audits to monitor for regressions.