SiteAud!tLint Support

robots.txt, AI crawlers and llms.txt

How SiteAud!tLint reads robots.txt, reports blocked URLs, checks which AI crawlers like GPTBot and ClaudeBot are allowed, and looks for llms.txt.

robots.txt tells crawlers which URLs they may fetch. SiteAud!tLint reads it at the start of every audit, uses it to find XML sitemaps, and checks it for answer engine readiness.

Blocked URLs

With Respect robots.txt on (the default), disallowed URLs are not fetched. They appear in URL health as Blocked and as the issue Blocked by robots.txt. If a blocked page should rank, remove or narrow the matching Disallow rule.

Untick Respect robots.txt only for sites you own, such as a staging copy that blocks everything.

Missing robots.txt

If /robots.txt returns an error, crawlers fall back to crawling everything and you lose the place to declare sitemaps. Publish one at the site root with a Sitemap: line:

User-agent: *
Disallow:

Sitemap: https://www.example.com/sitemap.xml

AI crawlers and answer engines

Improve > AEO / GEO shows which AI and answer engine crawlers your robots.txt allows. When some are disallowed, the site may not be used or cited in AI answers, and SiteAud!tLint reports AI crawlers blocked in robots.txt. Blocking can be a deliberate choice; the point is to decide it on purpose.

llms.txt

/llms.txt is an emerging convention: a short Markdown file giving language models a summary of the site and links to its most important pages. SiteAud!tLint checks for it and reports No llms.txt file as an info notice, which never lowers the score.

Applies to SiteAud!tLint 0.3 and later. Current version 0.7.1.