AI Crawlers and User Agents
AI companies use three kinds of bot: training crawlers that collect data for building models, search crawlers that build an index for AI answers, and user-triggered fetchers that retrieve a page when someone asks a question. Each has its own user agent and can be allowed or blocked separately.
  • Know the difference between training, search and user-triggered bots before you block anything.
  • Blocking a search crawler removes you from that platform's AI answers.
  • Google-Extended controls Gemini training and grounding, not Google Search or AI Overviews.
  • Most AI crawlers do not run JavaScript.
  • Check your CDN and firewall settings. Many block AI bots by default.

Deciding which AI bots can access your site is now a routine part of technical SEO. The right answer depends on your business: a publisher worried about content being used for training may choose differently from a brand that wants to be recommended as widely as possible. Whatever you decide, decide it on purpose.

Three types of AI bot

  • Training crawlers gather content to train future models. Blocking them keeps your content out of training data and has no immediate effect on citations.
  • Search or index crawlers build the index an AI product searches when answering. Blocking them makes you much less likely to be cited by that product.
  • User-triggered fetchers retrieve a specific page in real time because a user asked about it or shared the link. These act on behalf of a person, and some providers state that they may not follow robots.txt in the same way as automated crawlers.

The main user agents

OpenAI

  • GPTBot: training.
  • OAI-SearchBot: search index for ChatGPT search.
  • ChatGPT-User: user-triggered fetches.

Anthropic

  • ClaudeBot: training.
  • Claude-SearchBot: search index.
  • Claude-User: user-triggered fetches.

Perplexity

  • PerplexityBot: search index.
  • Perplexity-User: user-triggered fetches.

Google

  • Googlebot: crawls for Search, which includes AI Overviews and AI Mode. There is no separate crawler for those features.
  • Google-Extended: a robots.txt control token, not a crawler with its own user agent string. It governs whether your content is used to train Gemini models and for grounding in Gemini products. It does not affect Search rankings or inclusion in AI Overviews.

Microsoft

  • Bingbot: crawls for Bing, which grounds Copilot and supplies results to other assistants.

Others

  • Applebot and Applebot-Extended: Apple search features, and the opt-out token for AI training.
  • Meta-ExternalAgent: Meta's crawler for AI training and products.
  • Amazonbot: Amazon, including Alexa.
  • CCBot: Common Crawl, an open dataset used widely in model training.
  • Bytespider: ByteDance.
  • DuckAssistBot and MistralAI-User: answer features from DuckDuckGo and Mistral.

New agents appear regularly and names change. Check each provider's documentation for the current list and published IP ranges.

Example robots.txt

Allow AI search and user fetches, block training:

 User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: / 

If your goal is maximum visibility, the simplest policy is to allow them all. Being in training data shapes what models say about you when they are not searching.

How AI crawlers behave

  • No JavaScript rendering, in most cases. They read the HTML response.
  • Short timeouts. Slow pages get skipped.
  • Uneven crawl patterns. Some crawl heavily and send little traffic back. Server logs show the ratio.
  • Variable compliance. Major providers honour robots.txt for their automated crawlers. Less scrupulous bots ignore it or disguise themselves.

Beyond robots.txt

  • CDN and firewall rules: Cloudflare and similar services offer one-click blocking of AI bots, and Cloudflare has blocked AI crawlers by default for new sites since mid-2025. Check your configuration.
  • Bot verification: validate requests against published IP ranges, since user agent strings can be spoofed.
  • Rate limiting: returning 429 to aggressive bots is reasonable. Make sure it does not catch the ones you want.
  • Pay-per-crawl and licensing schemes: emerging options for publishers who want compensation for access.

Common mistakes

  • Blocking GPTBot and assuming that removes you from ChatGPT search. That is OAI-SearchBot's job.
  • Blocking Google-Extended and expecting to disappear from AI Overviews.
  • A security product silently returning 403 to all AI agents.
  • Allowing bots in robots.txt while serving them an empty JavaScript shell.
  • Copying a blocklist from a blog post without checking what each agent does.

How to test

  • Filter server or CDN logs by AI user agent and review request volumes and status codes.
  • Fetch a page with curl using an AI user agent string to see what comes back.
  • Ask an assistant to read a specific URL from your site and see whether it can.
  • Review bot analytics in your CDN dashboard.

Add your title here

This is a paragraph. Writing in paragraphs lets visitors find what they are looking for quickly and easily. Make sure the title suits the content of this text.

Contact Us Amy Time