AI Crawlers vs Traditional Search Crawlers

AI OVERVIEW

Blocking an AI training bot doesn't block that company's search bot or user fetcher. Each needs its own robots.txt rule.

For twenty years, "crawler" mostly meant Googlebot and Bingbot. Now your server logs are full of GPTBot, ClaudeBot, PerplexityBot, and a dozen others, and each one wants something different from your pages.

The biggest misunderstanding is treating "AI crawlers" as one thing. Most AI companies run several bots with different jobs, and a robots.txt rule for one doesn't touch the others.

Last reviewed: [add date]. Bot names and policies change often, so check each operator's documentation before changing your robots.txt.

Traditional Search Crawlers

Googlebot and Bingbot do one main job: fetch pages so they can be indexed and shown in search results.

  • They follow robots.txtRules are well documented and consistently honored.
  • They render JavaScriptGooglebot uses a recent version of Chromium.
  • They adapt crawl rateSpeed adjusts to how your server responds.
  • They can be verifiedOperators publish ways to confirm a request really came from them.

Blocking them takes you out of their search results. That's why a stray Disallow: / for Googlebot is one of the most damaging mistakes in technical SEO.

AI Crawlers Come in Three Types

The group a bot belongs to decides what blocking it actually does.

Type 1

Training crawlers

Collect public content that may train future models. Blocking affects future training only, not AI answers today.

Type 2

Search index crawlers

Build the index an AI product searches when answering. Block these and you're less likely to be found and cited.

Type 3

User-triggered fetchers

Visit a page in real time because a user asked. Some operators treat them differently under robots.txt.

OperatorTrainingSearch indexUser-triggered
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexity(none listed)PerplexityBotPerplexity-User
One AI company Training botDisallow: / Search botstill allowed User fetcherstill allowed Each name is a separate robots.txt token with its own rule.
Anthropic, for example, documents all three bots separately. Blocking ClaudeBot doesn't block Claude-SearchBot or Claude-User.

User-triggered fetchers are where policies differ most. Anthropic states that all its bots, including Claude-User, honor robots.txt. OpenAI and Perplexity have said user-initiated fetches may not follow robots.txt the same way automated crawling does. If keeping content out of AI tools completely matters, robots.txt alone may not be enough.

Tokens that aren't crawlers

  • Google-ExtendedControls whether content fetched by Google's normal crawlers can be used for Gemini training and grounding. It doesn't affect Google Search. AI Overviews are part of Search, so blocking Google-Extended doesn't remove you from them.
  • Applebot-ExtendedWorks the same way for Apple's AI training. Applebot itself still crawls for Siri and Spotlight.
  • CCBot and othersCommon Crawl's public dataset is widely used for training, and bots from Meta, ByteDance, Amazon, and others also visit. Not all document their behavior clearly.

What Actually Differs

Search crawlersAI crawlers
Main purposeBuild a search indexTraining, an AI search index, or live fetches for a user
Bots per operatorMostly one main botOften three, each controlled separately
JavaScriptGooglebot and Bingbot render itMany appear to read only raw HTML
Effect of blockingRemoved from search resultsDepends entirely on which bot you block
VerificationReverse DNS and published IP rangesPublished IP ranges from some operators, inconsistent across others
Crawl-delayIgnored by Googlebot, respected by BingSupported by some (Anthropic documents it), not all
The JavaScript row matters most

If main content only appears after client-side rendering, a crawler that reads raw HTML sees an empty shell. Content you want AI search tools to find belongs in the HTML your server sends. Our JavaScript SEO guide covers raw vs rendered HTML.

robots.txt for AI Crawlers

Each bot follows the most specific group that names it. A common setup allows AI search and live fetches but opts out of training:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml
  • A named group replaces the * groupOnce you add User-agent: GPTBot, rules under User-agent: * like Disallow: /admin/ no longer apply to GPTBot. Repeat them in each named group.
  • robots.txt is a request, not a lockWell-behaved bots follow it. Scrapers pretending to be GPTBot won't. Use authentication or server-level blocking for private content.
  • Blocking by IP can backfireAnthropic notes that blocking its IPs may stop its bots from reading your robots.txt at all, so an IP-based opt-out may not work reliably.
  • Old tokens lingerMany files still list deprecated names like anthropic-ai or Claude-Web. Harmless, but they don't control current bots.

Verifying Who's Really Crawling

User-agent strings are trivial to fake. Anyone can send User-Agent: Googlebot. Verify the important ones before acting on your logs:

Request IPFrom your access log
→
Reverse DNSgooglebot.com, google.com, or search.msn.com
→
Forward DNSHostname returns the same IP
→
VerifiedOr match published IP ranges
  • GooglebotReverse DNS should resolve to googlebot.com or google.com, then a forward lookup should return the same IP. Google also publishes its crawler IP ranges as JSON.
  • BingbotReverse DNS should resolve to search.msn.com. Bing Webmaster Tools also has a verification tool.
  • AI crawlersSome operators, including OpenAI, publish IP ranges. A "GPTBot" request from an IP not on OpenAI's list isn't GPTBot.

What Your Server Logs Can Tell You

A site crawl shows what bots could access. Server logs show what they actually requested. You need both.

  • Which bots visit at allMany sites find AI bots visit far more often than expected, or not at all.
  • Status codes per botMostly 403s or 429s means something is blocking it, often a firewall or CDN bot-protection rule.
  • What they requestImportant pages, or mostly old URLs, parameters, and 404s?
  • Whether they obey robots.txtRequests from a verified bot to a disallowed path are worth reporting to the operator.
  • LoadA Crawl-delay (for bots that support it) or temporary 429 and 503 responses can slow a heavy bot. Google advises against returning those to Googlebot for long periods.

Logs usually live with your host, CDN, or server, not in a crawler, so this part needs access to those systems.

A Quick Crawler Audit

  • Read your robots.txt as each bot would: which group applies, and what it allows.
  • Check blocks are intentional: leftover Disallow rules, deprecated tokens, named groups that dropped your * rules.
  • Check CDN and firewall bot settings, which can block AI search crawlers whatever robots.txt says.
  • Confirm key content is in the raw HTML.
  • Review a week of logs by verified bot, status code, and URL type.
  • Recheck after every robots.txt change, CDN rule, or new bot in an operator's lineup.

Checking AI Crawler Access With SiteAuditLint

The SiteAuditLint AEO / GEO audit checks nine major AI crawlers, including GPTBot, ClaudeBot, and PerplexityBot, against your robots.txt and shows which are allowed or blocked. It also detects llms.txt, a proposed file for guiding AI tools that major search engines haven't confirmed as a ranking signal, and flags content structure issues. It all rolls up into an Answer Engine Readiness score.

Access is only half the question. The AI visibility tracker runs your own prompts through ChatGPT, Gemini, Claude, and Perplexity and reports how often your pages are cited, so you can see whether allowing those crawlers is paying off.

[Add screenshot here: the AEO / GEO crawler access view from a real audit, showing allowed and blocked AI bots.]
Decide per bot, write each rule explicitly, verify who's really visiting, and keep the content you want found in the raw HTML.