The Trap Room
Please lower your voice. Some of them are still in the net.
- The site's robots.txt says, plainly:
Disallow: /trap/ - Every page hides a link to
/trap/that no human can see or click. - Anything that walks through it has ignored the sign and come in anyway.
The Descent
LIVEA cutaway of the Labyrinth, right now. Every bot inside hangs at the depth it has reached; when it opens a door, it drops another room.
In the net
Most recent catches · tagged on captureSpecimens in disguise
Each of these arrived announcing itself as a well-known crawler. Real Googlebot comes from Google's network, and its address resolves to a google.com or googlebot.com hostname. These did not. (We don't keep visitors' addresses, so here is the network they came from instead.)
Escaped & recaptured
The trap does not hold anything. It only remembers. These keep coming back.
The Labyrinth
Behind the staff door, a staircase goes down into a maze with no end. Every room has five doors, and every door leads deeper. Only bots that ignore robots.txt ever find it, and none has found the bottom, because there isn't one.
Add a disallowed path to your robots.txt, link to it invisibly, and log whatever arrives. Anything in that log ignored your rules.
User-agent: * Disallow: /trap/