Home›Academy›AI Search›AI crawlers

AI Search · Lesson 03 · Access

AI crawlers

Which bots collect content for AI systems, how training, search, and user-triggered crawlers differ, and how to control and monitor them.

06 AI Search11 min Read timeIntermediate Level

03 / Access

AI crawlers are bots that fetch web content for AI companies. Some collect training data, some build search indexes for AI answers, and some fetch a page live when a user asks a question. Each type has different consequences for visibility, so a blanket allow or block decision is rarely the right one.

This lesson identifies the main AI user agents, explains how robots.txt and firewalls control them, and shows how to monitor AI crawler activity.

1. Three types of AI crawler

TypePurposeEffect of blocking
Training crawlersCollect content to train future modelsContent is not used in future training. Little direct effect on citations
Search and index crawlersBuild an index used to cite sources in AI answersPages may stop appearing as cited sources
User-triggered fetchersFetch a page when a user asks the assistant to read or check itThe assistant cannot read the page on request

2. Common AI user agents

CompanyTrainingSearch or indexUser-triggered
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
PerplexityNot documented as a training crawlerPerplexityBotPerplexity-User
GoogleGoogle-Extended (a control token, not a separate crawler)GooglebotGoogle user-triggered fetchers
AppleApplebot-Extended (control token)Applebot
Common CrawlCCBot (dataset used by many AI projects)
Check the current documentation

AI platforms change crawler names, ranking systems, and citation formats often. The details in this lesson reflect common behaviour at the time of writing. Confirm user agent names and settings in each provider's official documentation before changing robots.txt or firewall rules.

Google-Extended controls whether content is used for Gemini training and grounding. It does not affect Google Search or AI Overviews, which use Googlebot. Blocking Googlebot to avoid AI Overviews would remove the site from Google Search entirely.

3. Controlling AI crawlers in robots.txt

Most documented AI crawlers respect robots.txt. A policy that blocks training but allows AI search could look like this:

# Allow AI search and user fetches
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

# Opt out of model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /

Test rules carefully. A broad User-agent: * disallow can block AI crawlers unintentionally. See the robots.txt lesson and the robots.txt and AI crawlers guide.

4. Firewalls, CDNs, and bot protection

Many sites block AI crawlers without knowing it. CDN bot management, web application firewalls, and hosting security features can return 403 or challenge pages to AI user agents, overriding a permissive robots.txt.

A hidden block

robots.txtUser-agent: PerplexityBot
Allow: /
The policy allows the crawler.
Server response403 Forbidden
Bot protection challenge
The firewall blocks it anyway.

Always test the actual response, not only the robots.txt file.

See fixing 403 errors blocking AI crawlers and using Cloudflare to monitor AI crawlers.

5. Verifying real crawlers

User agent strings are easy to fake. Scrapers often claim to be well-known bots. Most major providers publish IP ranges for their crawlers, so verify by IP before allowlisting and avoid granting special access on user agent alone.

6. llms.txt and markdown versions

llms.txt is a proposed file that lists a site's key content in markdown for language models. Adoption by major AI crawlers is limited and contested, so treat it as optional. Some AI agents also request markdown versions of pages. See the llms.txt debate and what text/markdown requests mean. SiteAuditLint reports llms.txt missing as informational, not critical.

7. Monitoring AI crawler activity

  1. Review server logs or CDN analyticsFilter requests by AI user agents.
  2. Verify IP rangesSeparate genuine crawlers from impersonators.
  3. Check response codesLook for 403, 429, and 5xx responses served to AI bots.
  4. Compare crawled URLsConfirm important pages are fetched, not only the homepage.
  5. Match with referralsConnect crawler activity with AI referral traffic.

8. Checking AI crawler access in SiteAuditLint

SiteAuditLint tests robots.txt rules for major AI user agents and flags AI crawlers blocked. Run a crawl with an AI user agent to confirm that firewalls return 200 rather than 403. The article on checking crawlability for AI bots covers manual tests.

9. Practical exercise: Write an AI crawler policy

AI crawler checklist

0 of 8 tasks completed

Key takeaways

  • AI crawlers serve training, search indexing, or user-triggered fetching.
  • Blocking training crawlers differs from blocking AI search crawlers.
  • Google-Extended does not control AI Overviews. Googlebot does.
  • Firewalls and CDNs often block AI crawlers despite robots.txt.
  • Verify crawler IPs because user agents can be faked.
  • Monitor logs to confirm AI crawlers reach important pages.

Knowledge check

1. What is the effect of blocking a search or index AI crawler?
2. robots.txt allows PerplexityBot but it receives 403. What is likely?
3. What does Google-Extended control?
4. Why verify crawler IP addresses?

Next, learn how AI systems choose and display sources in the AI citations lesson.