Vol. I · Fri, 2 Oct 2026Open 24 hours · Feeding: continuous · Please do not tap the glass
Plate XIX0xAEEA4B00UNREAD
Kingdom Automata›Phylum HTTP›Order Crawlers›Family AI Scrapers›Species CCBot
Specimen file · The AI Aviary

CCBot

Communis omnivorus

What is CCBot?

Common Crawl's crawler. It builds a free public archive of the web that many AI training sets are drawn from. It is operated by Common Crawl and identifies itself as CCBot.

Temperament: generous. Eats everything and shares it with everyone. Nobody asked.

Operator
Common Crawl
Conservation status
Has not read the rules
First seen
Not yet
Last seen
Not yet
Visits today
0
This week
0
All time
0
Caught in the trap
Never
How to block CCBot ↓
§ I.

Identification

Its full user-agent string, exactly as it arrives at the gate:

User-Agent
$ CCBot/2.0 (https://commoncrawl.org/faq/)

Common Crawl does not publish a way to verify CCBot's traffic. Treat its user agent as a claim, not proof: anyone can send it. Read Common Crawl's documentation.

§ II.

How to block CCBot

Add these two lines to the robots.txt file at the root of your site. Well-behaved crawlers read it before they crawl, so the change applies from CCBot's next visit. Nothing else on your site needs to change.

robots.txt
# Block CCBot from the whole site
User-agent: CCBot
Disallow: /

Or let it visit but keep it away from part of the site:

# Let it in, but keep it out of one room
User-agent: CCBot
Allow: /
Disallow: /members/
§ III.

Observed behaviour

Visits, last 30 dayspeak 1 / day
3 Sept 2026today
Diet · pages most often eaten

Nothing on record in the last 30 days.

Visiting hours
00h – 11h12h – 23h
robots.txt compliance
100.0%

Requested 0 disallowed pages out of 0 requests. Never read robots.txt.

Trap Room record
Clean

Has never followed the hidden link to /trap/. Either well trained or very lucky.

§ IV.

Where it comes from

The networks CCBot's visits came from:

Nothing on record in the last 30 days.

Networks and countries come from the visitor's IP address, looked up in a local copy of the DB-IP database. The addresses themselves are never stored.

§ V.

Questions site owners ask

Does CCBot respect robots.txt?

We can't say yet. It has not fetched robots.txt here, and it has not touched a disallowed page either.

Will blocking CCBot hurt my search rankings?

No. CCBot isn't used by Google or Bing to rank pages. Blocking it only affects whether Common Crawl collects your content.

How often does CCBot visit?

Here, about 0 requests a day over the last week. Visits to your site depend on its size, how often it changes, and how many links point to it.

Also in The AI Aviary