For twenty years, "crawler" mostly meant Googlebot and Bingbot. Now your server logs are full of GPTBot, ClaudeBot, PerplexityBot, and a dozen others, and each one wants something different from your pages.
The biggest misunderstanding is treating "AI crawlers" as one thing. Most AI companies run several bots with different jobs, and a robots.txt rule for one doesn't touch the others.
Last reviewed: [add date]. Bot names and policies change often, so check each operator's documentation before changing your robots.txt.
Traditional Search Crawlers
Googlebot and Bingbot do one main job: fetch pages so they can be indexed and shown in search results.
- They follow robots.txtRules are well documented and consistently honored.
- They render JavaScriptGooglebot uses a recent version of Chromium.
- They adapt crawl rateSpeed adjusts to how your server responds.
- They can be verifiedOperators publish ways to confirm a request really came from them.
Blocking them takes you out of their search results. That's why a stray Disallow: / for Googlebot is one of the most damaging mistakes in technical SEO.
AI Crawlers Come in Three Types
The group a bot belongs to decides what blocking it actually does.
Training crawlers
Collect public content that may train future models. Blocking affects future training only, not AI answers today.
Search index crawlers
Build the index an AI product searches when answering. Block these and you're less likely to be found and cited.
User-triggered fetchers
Visit a page in real time because a user asked. Some operators treat them differently under robots.txt.
| Operator | Training | Search index | User-triggered |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | (none listed) | PerplexityBot | Perplexity-User |
User-triggered fetchers are where policies differ most. Anthropic states that all its bots, including Claude-User, honor robots.txt. OpenAI and Perplexity have said user-initiated fetches may not follow robots.txt the same way automated crawling does. If keeping content out of AI tools completely matters, robots.txt alone may not be enough.
Tokens that aren't crawlers
- Google-ExtendedControls whether content fetched by Google's normal crawlers can be used for Gemini training and grounding. It doesn't affect Google Search. AI Overviews are part of Search, so blocking Google-Extended doesn't remove you from them.
- Applebot-ExtendedWorks the same way for Apple's AI training. Applebot itself still crawls for Siri and Spotlight.
- CCBot and othersCommon Crawl's public dataset is widely used for training, and bots from Meta, ByteDance, Amazon, and others also visit. Not all document their behavior clearly.
What Actually Differs
| Search crawlers | AI crawlers | |
|---|---|---|
| Main purpose | Build a search index | Training, an AI search index, or live fetches for a user |
| Bots per operator | Mostly one main bot | Often three, each controlled separately |
| JavaScript | Googlebot and Bingbot render it | Many appear to read only raw HTML |
| Effect of blocking | Removed from search results | Depends entirely on which bot you block |
| Verification | Reverse DNS and published IP ranges | Published IP ranges from some operators, inconsistent across others |
| Crawl-delay | Ignored by Googlebot, respected by Bing | Supported by some (Anthropic documents it), not all |
If main content only appears after client-side rendering, a crawler that reads raw HTML sees an empty shell. Content you want AI search tools to find belongs in the HTML your server sends. Our JavaScript SEO guide covers raw vs rendered HTML.
robots.txt for AI Crawlers
Each bot follows the most specific group that names it. A common setup allows AI search and live fetches but opts out of training:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
- A named group replaces the * groupOnce you add
User-agent: GPTBot, rules underUser-agent: *likeDisallow: /admin/no longer apply to GPTBot. Repeat them in each named group. - robots.txt is a request, not a lockWell-behaved bots follow it. Scrapers pretending to be GPTBot won't. Use authentication or server-level blocking for private content.
- Blocking by IP can backfireAnthropic notes that blocking its IPs may stop its bots from reading your robots.txt at all, so an IP-based opt-out may not work reliably.
- Old tokens lingerMany files still list deprecated names like
anthropic-aiorClaude-Web. Harmless, but they don't control current bots.
Verifying Who's Really Crawling
User-agent strings are trivial to fake. Anyone can send User-Agent: Googlebot. Verify the important ones before acting on your logs:
- GooglebotReverse DNS should resolve to
googlebot.comorgoogle.com, then a forward lookup should return the same IP. Google also publishes its crawler IP ranges as JSON. - BingbotReverse DNS should resolve to
search.msn.com. Bing Webmaster Tools also has a verification tool. - AI crawlersSome operators, including OpenAI, publish IP ranges. A "GPTBot" request from an IP not on OpenAI's list isn't GPTBot.
What Your Server Logs Can Tell You
A site crawl shows what bots could access. Server logs show what they actually requested. You need both.
- Which bots visit at allMany sites find AI bots visit far more often than expected, or not at all.
- Status codes per botMostly 403s or 429s means something is blocking it, often a firewall or CDN bot-protection rule.
- What they requestImportant pages, or mostly old URLs, parameters, and 404s?
- Whether they obey robots.txtRequests from a verified bot to a disallowed path are worth reporting to the operator.
- LoadA
Crawl-delay(for bots that support it) or temporary429and503responses can slow a heavy bot. Google advises against returning those to Googlebot for long periods.
Logs usually live with your host, CDN, or server, not in a crawler, so this part needs access to those systems.
A Quick Crawler Audit
- Read your robots.txt as each bot would: which group applies, and what it allows.
- Check blocks are intentional: leftover Disallow rules, deprecated tokens, named groups that dropped your * rules.
- Check CDN and firewall bot settings, which can block AI search crawlers whatever robots.txt says.
- Confirm key content is in the raw HTML.
- Review a week of logs by verified bot, status code, and URL type.
- Recheck after every robots.txt change, CDN rule, or new bot in an operator's lineup.
Checking AI Crawler Access With SiteAuditLint
The SiteAuditLint AEO / GEO audit checks nine major AI crawlers, including GPTBot, ClaudeBot, and PerplexityBot, against your robots.txt and shows which are allowed or blocked. It also detects llms.txt, a proposed file for guiding AI tools that major search engines haven't confirmed as a ranking signal, and flags content structure issues. It all rolls up into an Answer Engine Readiness score.
Access is only half the question. The AI visibility tracker runs your own prompts through ChatGPT, Gemini, Claude, and Perplexity and reports how often your pages are cited, so you can see whether allowing those crawlers is paying off.