Identification
Its full user-agent string, exactly as it arrives at the gate:
There is no operator to verify against. This name is what software calls itself when nobody gave it one.
How to block curl
curl does not read robots.txt, so a polite sign is wasted on it. Refuse it at your web server or firewall instead. User agents are easy to fake, so pair this with rate limiting.
# robots.txt will not stop curl. Block it at the server.
# nginx
if ($http_user_agent ~* "curl") {
return 403;
}The same thing on Apache:
# Apache (.htaccess)
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} curl [NC]
RewriteRule .* - [F,L]Observed behaviour
Most active around 16:00. Office hours, like a professional.
Requested 1 disallowed pages out of 30 requests. Read robots.txt 3 times.
Has never followed the hidden link to /trap/. Either well trained or very lucky.
Where it comes from
Scripts and scanners run from wherever their owners rent a server. These are the networks behind the visits on file:
Networks and countries come from the visitor's IP address, looked up in a local copy of the DB-IP database. The addresses themselves are never stored.
Keeper's field notes
Questions site owners ask
Does curl respect robots.txt?
Mostly. It reads robots.txt, but it has fetched a disallowed page 1 times out of 30 requests observed here.
Will blocking curl hurt my search rankings?
No. Nothing respectable will miss it.
How often does curl visit?
Here, about 4 requests a day over the last week. Visits to your site depend on its size, how often it changes, and how many links point to it.