Vol. I · Fri, 2 Oct 2026Open 24 hours · Feeding: continuous · Please do not tap the glass
Plate XII0x6A85CEA7UNREAD
Kingdom Automata›Phylum HTTP›Order Crawlers›Family Search Engines›Species archive.org_bot
Specimen file · Search Engine Savanna

archive.org_bot

Memoria aeterna

What is archive.org_bot?

The Internet Archive's crawler. It saves copies of pages for the Wayback Machine. It is operated by Internet Archive and identifies itself as archive.org_bot.

Temperament: sentimental. Keeps a copy of everything, just in case. Has photographs of this page as a child.

Operator
Internet Archive
Conservation status
Has not read the rules
First seen
1 Oct 2026
Last seen
22 h ago
Visits today
2
This week
2
All time
2
Caught in the trap
Never
§ I.

Identification

Its full user-agent string, exactly as it arrives at the gate:

User-Agent
$ Mozilla/5.0 (compatible; archive.org_bot +http://archive.org/details/archive.org_bot)

Internet Archive does not publish a way to verify archive.org_bot's traffic. Treat its user agent as a claim, not proof: anyone can send it. Read Internet Archive's documentation.

§ II.

How to block archive.org_bot

Add these two lines to the robots.txt file at the root of your site. Well-behaved crawlers read it before they crawl, so the change applies from archive.org_bot's next visit. Nothing else on your site needs to change.

robots.txt
# Block archive.org_bot from the whole site
User-agent: archive.org_bot
Disallow: /

Or let it visit but keep it away from part of the site:

# Let it in, but keep it out of one room
User-agent: archive.org_bot
Allow: /
Disallow: /members/
§ III.

Observed behaviour

Visits, last 30 dayspeak 2 / day
3 Sept 2026today
Diet · pages most often eaten
/og/home.png1
/1
Visiting hours
00h – 11h12h – 23h

Most active around 21:00. After dark, like a burglar.

robots.txt compliance
100.0%

Requested 0 disallowed pages out of 2 requests. Never read robots.txt.

Trap Room record
Clean

Has never followed the hidden link to /trap/. Either well trained or very lucky.

§ IV.

Where it comes from

The networks archive.org_bot's visits came from:

Networks · 1 on file, last 30 days
Internet Archive Canada AS399784100%
Countries
Canada CA100%

Networks and countries come from the visitor's IP address, looked up in a local copy of the DB-IP database. The addresses themselves are never stored.

§ V.

Keeper's field notes

1 Oct 202621:31. Arrived without announcement. Consumed /og/home.png and /. Departed.
2 Oct 2026Visited 2 times in a single day. Keeper unable to establish why.
1 Oct 2026First recorded at the gate. Entered in the register as Memoria aeterna.
§ VI.

Questions site owners ask

Does archive.org_bot respect robots.txt?

We can't say yet. It has not fetched robots.txt here, and it has not touched a disallowed page either.

Will blocking archive.org_bot hurt my search rankings?

Yes, for Internet Archive's search engine. Blocked pages can drop out of its results. Other search engines are unaffected.

How often does archive.org_bot visit?

Here, about 0 requests a day over the last week. Visits to your site depend on its size, how often it changes, and how many links point to it.

Also in Search Engine Savanna