Add the training crawlers to your robots.txt. It takes a minute, costs nothing in search, and the well-known AI companies respect it.
Three kinds of AI bot
"AI crawler" covers three different jobs. They have different names, so you can choose which to keep out.
Training crawlers
They collect pages that may be used to train AI models. Blocking them is the usual answer to "I don’t want my writing in an AI model". Blocking them costs you nothing in search results.
| Bot | Operator | robots.txt name | Seen here this week |
|---|---|---|---|
| GPTBot | OpenAI | GPTBot | — |
| ClaudeBot | Anthropic | ClaudeBot | 38 |
| CCBot | Common Crawl | CCBot | — |
| Bytespider | ByteDance | Bytespider | — |
| Meta-ExternalAgent | Meta | meta-externalagent | — |
| cohere-ai | Cohere | cohere-ai, cohere-training-data-crawler | — |
| AI2Bot | Allen Institute for AI | AI2Bot | — |
| Diffbot | Diffbot | Diffbot | — |
| ImagesiftBot | Hive | ImagesiftBot | — |
| Google-Extended | Google-Extended | rule only | |
| Applebot-Extended | Apple | Applebot-Extended | rule only |
AI search crawlers
They index pages so AI assistants can find and cite them when answering questions. Blocking them can keep you out of those assistants’ answers, and out of the traffic their citations send.
| Bot | Operator | robots.txt name | Seen here this week |
|---|---|---|---|
| OAI-SearchBot | OpenAI | OAI-SearchBot | — |
| Claude-SearchBot | Anthropic | Claude-SearchBot | — |
| PerplexityBot | Perplexity | PerplexityBot | — |
| YouBot | You.com | YouBot | — |
| DuckAssistBot | DuckDuckGo | DuckAssistBot | — |
| Amazonbot | Amazon | Amazonbot | — |
User-triggered fetchers
They fetch one page because one person asked an assistant about it. They are closer to a browser than a crawler. Several operators say robots.txt may not apply to them, since a person asked.
| Bot | Operator | robots.txt name | Seen here this week |
|---|---|---|---|
| ChatGPT-User | OpenAI | ChatGPT-User | — |
| Claude-User | Anthropic | Claude-User | — |
| Perplexity-User | Perplexity | Perplexity-User | — |
| MistralAI-User | Mistral AI | MistralAI-User | — |
Ready-made rules
Copy one of these into the robots.txt file at the root of your site, or open it in the generator to adjust it.
Keep AI training out, stay in search and AI answers
The most common choice. Search engines and AI search crawlers stay welcome.
User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: Bytespider User-agent: meta-externalagent User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: AI2Bot User-agent: Diffbot Disallow: /
Keep all AI out
Training, AI search and user-triggered fetchers. Ordinary search engines are unaffected.
User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot User-agent: Bytespider User-agent: meta-externalagent User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: AI2Bot User-agent: Diffbot User-agent: ImagesiftBot User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: YouBot User-agent: DuckAssistBot User-agent: Amazonbot User-agent: ChatGPT-User User-agent: Claude-User User-agent: Perplexity-User User-agent: MistralAI-User User-agent: Google-Extended User-agent: Applebot-Extended Disallow: /
Keep AI out of one part of your site
Replace / with a folder. Everything else stays open.
User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot User-agent: Google-Extended Disallow: /members/ Disallow: /drafts/
Google-Extended and Applebot-Extended
These two are not crawlers. They are names you can use in robots.txt to tell Google and Apple not to use your pages for their AI models. Googlebot and Applebot read the rule and carry on crawling for search, so blocking them keeps you in Google Search, Siri and Spotlight while opting out of Gemini and Apple Intelligence training.
Because they never visit, a server rule cannot block them. Only robots.txt can.
What robots.txt can't do
robots.txt is a request, not a lock. It works because the big companies choose to respect it. Three things it cannot stop:
- Bots that ignore it. In the last 30 days, 28 visitors walked through a door this site's robots.txt marks as off-limits. The Trap Room lists them.
- Impostors. Anyone can call their scraper GPTBot. Real crawlers come from their operators' networks, and several publish their IP ranges. This zoo checks every visit; failures are filed as impostors, like GPTBot (impostor).
- Data already collected. Blocking works from the next visit. It does not remove pages that were fetched before.
Going further: blocking at the server
For bots that ignore robots.txt, refuse them at your web server. A rule like this in nginx returns "403 Forbidden" to any listed name:
map $http_user_agent $ai_bot {
default 0;
~*(GPTBot|ClaudeBot|CCBot|Bytespider|meta-externalagent|cohere|AI2Bot|Diffbot) 1;
}
server {
# ...
if ($ai_bot) { return 403; }
}User agents are easy to fake, so a scraper can simply change its name. Pair this with rate limiting. If your site sits behind Cloudflare, its dashboard has a one-click switch to block known AI crawlers. (The zoo leaves it off, for obvious reasons.)
Signals that aren't rules
- llms.txt is a plain-text map of your site written for AI models. It invites them in rather than keeping them out. The zoo's llms.txt.
- ai.txt (from Spawning) states what AI may use your site for. Few crawlers read it yet.
- TDMRep is a W3C community standard for reserving text-and-data-mining rights, which matters under EU copyright rules. It uses a
/.well-known/tdmrep.jsonfile or atdm-reservationheader. - "noai" meta tags are not a standard. A few services honour them; most crawlers do not.
The Reading Room shows which bots actually open these files.
Check that it worked
Open https://your-site/robots.txt in a browser to confirm the file is live. Then watch your server logs for the names you blocked: well-behaved bots will fetch robots.txt and stop requesting your pages, usually within a day. Any that keep coming are ignoring the rule, and belong in a server block.
Questions
Will blocking GPTBot remove my site from ChatGPT?
Not from ChatGPT search. GPTBot collects training data; ChatGPT search uses OAI-SearchBot, which has its own name. Block both to keep out entirely.
Does blocking Google-Extended affect Google Search?
No. Google-Extended only controls whether Google may use your pages for its Gemini models. Googlebot still crawls and ranks your site as before.
Do AI crawlers actually obey robots.txt?
The big, named ones say they do, and most behave. Some do not, and anyone can send a fake name. The zoo’s Trap Room keeps a live record of who breaks the rules.
How long does a robots.txt change take to work?
Crawlers re-read robots.txt from time to time, usually within a day. Google may keep a cached copy for up to 24 hours.
Can I block AI bots from just part of my site?
Yes. Use Disallow with a path, for example Disallow: /blog/, instead of Disallow: / for the whole site.
Tick the bots you want to keep out, or start from a preset.