03 / Access
AI crawlers are bots that fetch web content for AI companies. Some collect training data, some build search indexes for AI answers, and some fetch a page live when a user asks a question. Each type has different consequences for visibility, so a blanket allow or block decision is rarely the right one.
This lesson identifies the main AI user agents, explains how robots.txt and firewalls control them, and shows how to monitor AI crawler activity.
1. Three types of AI crawler
| Type | Purpose | Effect of blocking |
|---|---|---|
| Training crawlers | Collect content to train future models | Content is not used in future training. Little direct effect on citations |
| Search and index crawlers | Build an index used to cite sources in AI answers | Pages may stop appearing as cited sources |
| User-triggered fetchers | Fetch a page when a user asks the assistant to read or check it | The assistant cannot read the page on request |
2. Common AI user agents
| Company | Training | Search or index | User-triggered |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | Not documented as a training crawler | PerplexityBot | Perplexity-User |
| Google-Extended (a control token, not a separate crawler) | Googlebot | Google user-triggered fetchers | |
| Apple | Applebot-Extended (control token) | Applebot | |
| Common Crawl | CCBot (dataset used by many AI projects) |
AI platforms change crawler names, ranking systems, and citation formats often. The details in this lesson reflect common behaviour at the time of writing. Confirm user agent names and settings in each provider's official documentation before changing robots.txt or firewall rules.
Google-Extended controls whether content is used for Gemini training and grounding. It does not affect Google Search or AI Overviews, which use Googlebot. Blocking Googlebot to avoid AI Overviews would remove the site from Google Search entirely.
3. Controlling AI crawlers in robots.txt
Most documented AI crawlers respect robots.txt. A policy that blocks training but allows AI search could look like this:
# Allow AI search and user fetches
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
# Opt out of model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /
Test rules carefully. A broad User-agent: * disallow can block AI crawlers unintentionally. See the robots.txt lesson and the robots.txt and AI crawlers guide.
4. Firewalls, CDNs, and bot protection
Many sites block AI crawlers without knowing it. CDN bot management, web application firewalls, and hosting security features can return 403 or challenge pages to AI user agents, overriding a permissive robots.txt.
A hidden block
Allow: /The policy allows the crawler.
Bot protection challengeThe firewall blocks it anyway.
Always test the actual response, not only the robots.txt file.
See fixing 403 errors blocking AI crawlers and using Cloudflare to monitor AI crawlers.
5. Verifying real crawlers
User agent strings are easy to fake. Scrapers often claim to be well-known bots. Most major providers publish IP ranges for their crawlers, so verify by IP before allowlisting and avoid granting special access on user agent alone.
6. llms.txt and markdown versions
llms.txt is a proposed file that lists a site's key content in markdown for language models. Adoption by major AI crawlers is limited and contested, so treat it as optional. Some AI agents also request markdown versions of pages. See the llms.txt debate and what text/markdown requests mean. SiteAuditLint reports llms.txt missing as informational, not critical.
7. Monitoring AI crawler activity
- Review server logs or CDN analyticsFilter requests by AI user agents.
- Verify IP rangesSeparate genuine crawlers from impersonators.
- Check response codesLook for 403, 429, and 5xx responses served to AI bots.
- Compare crawled URLsConfirm important pages are fetched, not only the homepage.
- Match with referralsConnect crawler activity with AI referral traffic.
8. Checking AI crawler access in SiteAuditLint
SiteAuditLint tests robots.txt rules for major AI user agents and flags AI crawlers blocked. Run a crawl with an AI user agent to confirm that firewalls return 200 rather than 403. The article on checking crawlability for AI bots covers manual tests.
9. Practical exercise: Write an AI crawler policy
AI crawler checklist
0 of 8 tasks completed
Key takeaways
- AI crawlers serve training, search indexing, or user-triggered fetching.
- Blocking training crawlers differs from blocking AI search crawlers.
- Google-Extended does not control AI Overviews. Googlebot does.
- Firewalls and CDNs often block AI crawlers despite robots.txt.
- Verify crawler IPs because user agents can be faked.
- Monitor logs to confirm AI crawlers reach important pages.
Knowledge check
Next, learn how AI systems choose and display sources in the AI citations lesson.