Identification
Its full user-agent string, exactly as it arrives at the gate:
This is not a real species. It is a costume: requests that use Googlebot's name but fail reverse-DNS or IP-range checks against Google's network.
How to block Googlebot (impostor)
Googlebot (impostor) does not read robots.txt, so a polite sign is wasted on it. Refuse it at your web server or firewall instead. User agents are easy to fake, so pair this with rate limiting.
# robots.txt will not stop Googlebot (impostor). Block it at the server.
# nginx
if ($http_user_agent ~* "Googlebot") {
return 403;
}The same thing on Apache:
# Apache (.htaccess)
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} Googlebot [NC]
RewriteRule .* - [F,L]Observed behaviour
Most active around 16:00. Office hours, like a professional.
Requested 0 disallowed pages out of 23 requests. Read robots.txt 3 times.
Has never followed the hidden link to /trap/. Either well trained or very lucky.
Where it comes from
Claims to be Googlebot. Real Googlebot traffic comes from Google's own network; these visits came from:
Networks and countries come from the visitor's IP address, looked up in a local copy of the DB-IP database. The addresses themselves are never stored.
Keeper's field notes
Questions site owners ask
Does Googlebot (impostor) respect robots.txt?
Yes. In 23 requests observed here it has read robots.txt and never fetched a disallowed page.
Will blocking Googlebot (impostor) hurt my search rankings?
No. Nothing respectable will miss it.
How often does Googlebot (impostor) visit?
Here, about 3 requests a day over the last week. Visits to your site depend on its size, how often it changes, and how many links point to it.