Vol. I · Fri, 2 Oct 2026Open 24 hours · Feeding: continuous · Please do not tap the glass
Plate XIII0x05FD0E0BUNREAD
Kingdom Automata›Phylum HTTP›Order Crawlers›Family AI Scrapers›Species GPTBot
Specimen file · The AI Aviary

GPTBot

Scraperus openaii

What is GPTBot?

OpenAI's web crawler. It collects publicly available pages that may be used to train OpenAI's models. It is operated by OpenAI and identifies itself as GPTBot.

Temperament: hungry. Visits at 3 a.m. Reads everything. Never says thank you.

Operator
OpenAI
Conservation status
Has not read the rules
First seen
Not yet
Last seen
Not yet
Visits today
0
This week
0
All time
0
Caught in the trap
Never
How to block GPTBot ↓
§ I.

Identification

Its full user-agent string, exactly as it arrives at the gate:

User-Agent
$ Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot

OpenAI publishes the IP ranges GPTBot uses. A request with this user agent from any other address is not GPTBot. The zoo checks every visit against that list. Read OpenAI's documentation.

§ II.

How to block GPTBot

Add these two lines to the robots.txt file at the root of your site. Well-behaved crawlers read it before they crawl, so the change applies from GPTBot's next visit. Nothing else on your site needs to change.

robots.txt
# Block GPTBot from the whole site
User-agent: GPTBot
Disallow: /

Blocking GPTBot does not remove you from ChatGPT search. That is OAI-SearchBot, which has its own token.

Or let it visit but keep it away from part of the site:

# Let it in, but keep it out of one room
User-agent: GPTBot
Allow: /
Disallow: /members/
§ III.

Observed behaviour

Visits, last 30 dayspeak 1 / day
3 Sept 2026today
Diet · pages most often eaten

Nothing on record in the last 30 days.

Visiting hours
00h – 11h12h – 23h
robots.txt compliance
100.0%

Requested 0 disallowed pages out of 0 requests. Never read robots.txt.

Trap Room record
Clean

Has never followed the hidden link to /trap/. Either well trained or very lucky.

§ IV.

Where it comes from

Requests that use GPTBot's name but fail OpenAI's network check are filed separately, as impostors. These are the visits that passed, or that could not be checked at the time:

Nothing on record in the last 30 days.

Networks and countries come from the visitor's IP address, looked up in a local copy of the DB-IP database. The addresses themselves are never stored.

§ V.

Questions site owners ask

Does GPTBot respect robots.txt?

We can't say yet. It has not fetched robots.txt here, and it has not touched a disallowed page either.

Will blocking GPTBot hurt my search rankings?

No. GPTBot isn't used by Google or Bing to rank pages. Blocking it only affects whether OpenAI collects your content.

How often does GPTBot visit?

Here, about 0 requests a day over the last week. Visits to your site depend on its size, how often it changes, and how many links point to it.

Also in The AI Aviary