How it works What we check Pricing
Guides
AEO GEO AI Search Visibility SEO Audit AI Citations llms.txt
Free tools
AI Visibility Check llms.txt Generator Robots Checker FAQ Schema
Learn
Blog Glossary FAQ AEO checklist About Sign in
Guides/llms txt/ai crawler list

AI Crawler User Agents List: Master Reference Table (2026)

In short

The main AI crawler user-agents in 2026 span three jobs: training (GPTBot, ClaudeBot, Google-Extended, CCBot), search indexing (OAI-SearchBot, Claude-SearchBot, PerplexityBot), and live user fetches (ChatGPT-User, Claude-User, Perplexity-User). Each is controlled independently in robots.txt.

01

How to read this list

Every AI operator runs more than one bot, and each bot does a different job. The single most useful thing to know is which job a user-agent does, because that determines what happens when you block it. Get this wrong and you either leak content you meant to protect or vanish from AI answers you wanted to appear in.

Three jobs:

  • Training — feeds long-term model training. Blocking = training opt-out, minimal visibility impact.

  • Search index — indexes pages to retrieve and cite in AI answers. Blocking = lost citations.

  • Live fetch — grabs a page because a user asked right now. Blocking = you miss real-time look-ups.

02

Master reference table

User-agent token Operator Job What blocking it does
GPTBot OpenAI Training Excludes content from OpenAI model training [1]
OAI-SearchBot OpenAI Search index Removes you from ChatGPT search citations
ChatGPT-User OpenAI Live fetch Stops live page fetches from ChatGPT / custom GPTs
ClaudeBot Anthropic Training Excludes content from Anthropic model training [2]
Claude-SearchBot Anthropic Search index Removes you from Claude search citations
Claude-User Anthropic Live fetch Stops live fetches when a user asks Claude
Google-Extended Google Training (signal) Opts out of Gemini/Vertex training; does not affect Search or AI Overviews [3]
Googlebot Google Search + AI Overviews Removes you from Google Search entirely — avoid
PerplexityBot Perplexity Search index Removes you from Perplexity's index (compliance caveat below)
Perplexity-User Perplexity Live fetch Stops user-triggered fetches from Perplexity
Meta-ExternalAgent Meta Training + fetch Excludes content from Meta AI training and on-demand fetch [4]
Bytespider ByteDance Training Excludes content from ByteDance model training (Doubao)
Amazonbot Amazon Search / assistant Removes content from Alexa and Amazon answer experiences
Applebot-Extended Apple Training (signal only) Opts out of Apple Intelligence training; does not crawl, does not affect Siri/Spotlight
CCBot Common Crawl Training (indirect) Excludes you from Common Crawl, which feeds many third-party training sets
03

Key distinctions people get wrong

GPTBot vs OAI-SearchBot. GPTBot is training. OAI-SearchBot is ChatGPT search. Independent. Block GPTBot and you're still in ChatGPT search. [1]

ClaudeBot vs Claude-SearchBot. ClaudeBot is training. Claude-SearchBot is search indexing. Claude-User is live user fetch. All independent. [2]

Google-Extended is not a crawler. It is a control token that governs whether Googlebot-crawled content is used for Gemini training. The crawler is still Googlebot. Blocking Google-Extended does not remove you from Search or from AI Overviews. [3]

Applebot-Extended does not crawl either. Same pattern: the live crawler is Applebot; Applebot-Extended is only the training opt-out signal. [4]

CCBot is Common Crawl, not any one AI company. Its archive is used as raw material by many model builders, so blocking CCBot has a broad, indirect training effect.

04

The Perplexity compliance note

PerplexityBot and Perplexity-User are the declared agents, and Perplexity states it respects robots.txt. However, Cloudflare documented cases where content was still accessed after those agents were disallowed, using an undeclared browser-like user-agent and rotating IPs. If you must guarantee a block, enforce it at the CDN/WAF layer rather than relying on robots.txt alone. [5]

05

Quick grouping for policy design

```txt

GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider,

Meta-ExternalAgent, Applebot-Extended

OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amazonbot

ChatGPT-User, Claude-User, Perplexity-User

06

FAQ

What is the difference between GPTBot and ChatGPT-User?

GPTBot crawls for training. ChatGPT-User fetches a page live when a user's question needs it. Different jobs, controlled separately.

Is Google-Extended a crawler?

No. It's a control token deciding whether Googlebot data trains Gemini. The crawler is Googlebot.

Which AI crawlers should I allow for visibility?

Keep the search-index and live-fetch bots open: OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User, Claude-User, Perplexity-User.

Does blocking CCBot block ChatGPT?

Not directly. CCBot is Common Crawl. Blocking it removes you from an archive many models train on, but OpenAI's own bots are separate.

Do these user-agents change?

Yes. Operators add and rename bots. Anthropic and OpenAI have both expanded their bot rosters recently. Re-check periodically.

How do I block one of these?

Add a User-agent: block with Disallow: / in robots.txt. For hard enforcement, mirror it in your CDN/WAF.

Keep this list current for your site

User-agents change, and your robots.txt drifts. SEO AEO Specialist checks your site against the live AI crawler roster and tells you exactly which bots you allow, block, or accidentally miss. Free at €0 for 50 pages per project, €9 one-time report, €19/mo per domain. Back to /llms-txt.

Sources: developers.openai.com, support.claude.com, muratulusoy.de, momenticmarketing.com, blog.cloudflare.com

How to Allow AI Crawlers: robots.txt + CDN/WAF Allowlist Guide

See where you stand

SEO AEO Specialist runs a free AI-visibility audit and hands you the exact fixes. One-off report €9, unlimited €19/mo.

Run your free audit →