That gap matters because bots are now the majority of requests. Imperva's 2026 Bad Bot Report found that automated traffic made up 53% of the web traffic it observed in 2025, and that bad bots alone made up 40%. So the useful question at your login form is rarely "is this a bot?" It's "which kind of bot, and does it get through?"
The kinds are worth naming: verified crawlers, unverified declared bots, impersonators, AI agents and user-triggered fetchers, scrapers, credential stuffers and scanners. Below we show what one lookup returned for five IPs, one per bot class, then point you to the guide that covers each job in depth: verifying crawlers by IP, reverse DNS and published ranges, protecting signups and checkout, and rate-limiting without locking out real users.
How to detect bots by IP address: six signals
To detect bots by IP address, look up the source IP and read six signals: the network type, reverse DNS, published crawler ranges, anonymity flags, scanner flags, and how many requests its prefix sends you. The IP is a better starting point than the User-Agent for one plain reason. A header is text the client chooses, but an HTTP request needs a completed TCP handshake, so the source IP can't simply be forged. The real workaround is renting or proxying someone else's IP, which the sections below cover.
We ran one anonymous GeoIPHub API lookup per bot class on 2026-10-07:
curl -s https://api.geoiphub.com/v1/lookup/66.249.66.1
Read down the score column and the main lesson shows. In GeoIPHub lookups on 2026-10-07, only the Tor exit was blocked (fraud_score 100). Googlebot scored 0, the Censys scanner 0, the iCloud Private Relay egress 9 and the DigitalOcean server 15, all allow. The score answers "how risky is this IP?", not "is this a bot?" A datacenter server at 15 and a scanner at 0 are both automation, and neither is risky on its own. List matches are the most reliable signals here; network-type flags are inferences, which is why they move the score only a little. These are five single lookups, not a benchmark, and values move as ranges change.
Here is what each signal tells you, and where the detail lives:
- Network type (
asn.asn_type,asn.user_type,detection.is_hosting). The ASN, the network that announces the IP, tells you whose network it is. A datacenter (hosting) IP, the API'sdetection.is_hosting, is the first filter, because most bots run in datacenters and most people don't. It's a weak verdict on its own, which is why DigitalOcean scored 15. Our explainer on what a datacenter IP address is covers how hosting is detected and when it's legitimate. - Reverse DNS (
network.ptr_record,fcrdns_valid). A forward-confirmed hostname can name the operator. The Tor exit's PTR ends intor-exit.artikel10.org, which firedrdns_tor. One limit of ours: the API stores no PTR for the Googlebot or DigitalOcean IPs, thoughdig -xreturns live names for both, sono_ptr_datacenterfires on them. Run the DNS check yourself when the hostname matters. - Published crawler ranges (
threat.crawler). Matching an IP to an operator's own range file is the cleanest proof a bot is who it says. GeoIPHub loads 9 official range feeds from Google, Bing, Apple, DuckDuckGo, OpenAI and Common Crawl, refreshed every 2 days. - Tor, VPN, proxy and relay flags (
detection.is_tor,is_vpn,is_proxy,is_relay). These say the real client is hidden behind someone else's network. See how to detect Tor traffic by IP and how VPN and proxy detection works. - Scanner flags (
threat.is_scanner). Internet-wide scanners probe every address they can reach. - Rate per prefix. This one is yours, not ours: count requests per /64 on IPv6, as the rate-limiting section below explains.
The scanner result surprised us, so here it is trimmed from the real response:
{
"ip": "162.142.125.10",
"asn": {
"asn": 398324,
"org": "CENSYS-ARIN-01 - Censys, Inc.",
"asn_type": "unknown"
},
"detection": { "is_hosting": false, "is_proxy": false, "is_vpn": false },
"threat": {
"is_scanner": true,
"threat_types": ["scanner"],
"crawler": { "verified": false, "name": "none" }
},
"scoring": {
"fraud_score": 0,
"recommended_action": "allow",
"detection_methods": []
}
}
The scanner flag is set, but it adds nothing to the score, and Censys's network isn't classified as hosting. If you rule only on recommended_action, scanner traffic passes. Read threat.is_scanner directly if you want to drop it from analytics or block it at the firewall.
Bot IP address check: is this crawler who it claims to be?
A bot IP address check confirms a declared crawler by matching its IP to the operator's published range file, or to a forward-confirmed reverse DNS name, never by trusting the User-Agent. Any client can send Googlebot in that header. Only the operator controls its IP ranges.
Our API sees only the IP, never your request headers, so your code makes the comparison: claimed name against threat.crawler.verified and threat.crawler.name. In our lookups, 66.249.66.1 returned verified: true, name: "Googlebot", verification_method: "ip_range_list". The DigitalOcean server returned name: "none", so a Googlebot claim from it is unverified. That means "not verified as that bot", not proof of a fake.
Our 9 feeds don't include ClaudeBot or PerplexityBot, so verify those with the operator's own method. Bots that sign their requests with Web Bot Auth can be checked by signature instead of by IP.
The detail lives in three posts. The crawler verification guide covers reverse DNS, CIDR files and Web Bot Auth signatures. How to verify Googlebot and spot fake Googlebots covers Google's three crawler groups and their range files. GPTBot vs OAI-SearchBot vs ChatGPT-User covers which OpenAI bot to block and how to check each one.
What IP-based bot detection can't tell you
IP-based bot detection can't tell you whether a home IP is being rented out as a residential proxy, and GeoIPHub does not detect residential proxies. A residential proxy network routes a bot's request through someone's home connection. At the IP layer that request looks like the household's own traffic: a consumer ISP, a real city, no hosting flag.
Our is_residential_proxy field is a coarse ASN-based inference, not a detection, so treat it as no signal. Rotation makes it worse. A pool can hand each request a different home IP, so per-IP counters never climb.
Botnets share the same blind spot. Infected home devices send bot traffic from ordinary residential IPs, and botnet feeds cover known infrastructure such as command-and-control servers, not every infected device. Bots get around IP detection in four common ways: rotating addresses, renting residential proxies, cycling through an IPv6 /64, and routing through mobile proxies that share carrier IPs with real phones.
- Residential proxies. Not detected. Catch them with behavior: velocity per account, device and session mismatches.
- The User-Agent and other headers. The API never sees them. Compare the claim in your own code.
- Who is typing. A headless browser and a person on the same office IP look identical at this layer.
- Ranges that changed in the last 2 days. Crawler and VPN feeds refresh every 2 days (the Tor exit list refreshes daily), so a brand-new range can be unverified for a short while.
- Crawlers outside our 9 feeds. No verified identity, whatever they claim.
- Scanner risk.
is_scanneris set but, in our lookup, doesn't raisefraud_score.
The layers that cover these gaps sit beside the IP: behavior and velocity, device fingerprinting, TLS fingerprints such as JA3 and JA4, challenges, and Web Bot Auth for bots that declare themselves. For the residential side, read residential proxy fraud detection for what a lookup can and can't show, and why so many residential proxies show up in your traffic for why there are so many.
How to detect bot traffic in signups, logins and checkout
To detect bot traffic at signups, logins and checkout, run the IP lookup at the moment of the action, then combine it with what your app already knows: attempts per account, per prefix and per device. The IP narrows the field. Your own counters make the call.
Each endpoint weighs the signals differently:
- Logins. Credential stuffing spreads attempts across many IPs, often residential ones the IP layer can't flag. Tor and hosting flags still catch the lazy share. Our account takeover and credential stuffing guide covers the layered playbook.
- Checkout. A datacenter IP at checkout may be a legitimate AI shopping agent, so hosting is a weighted input rather than a block: the DigitalOcean server scored 15,
allow. See how to stop AI shopping agent fraud at checkout. - Ad clicks. Hosting and Tor flags mark clicks worth keeping out of your conversion data. Our guide to detecting click fraud by IP address shows the scoring.
- Signups. The same bands apply, plus a check for many accounts from one prefix.
Use the score bands as the default action: allow up to 25, review up to 50, step_up up to 75, block above 75. Add your own rules for the flags the score leaves alone, like is_scanner.
How to rate-limit bot IPs without blocking real users
Rate-limit bot IPs by the unit that maps to one actor: the /64 on IPv6, and the single address on IPv4, adjusted for shared exits like CGNAT and iCloud Private Relay. Count the wrong unit and you either miss the bot or punish a crowd.
On IPv6, one subscriber gets a whole /64 and can rotate through it for free, so a per-address counter never trips. Our guide to rate-limiting and blocking IPv6 by /64 prefix shows the key function and when to widen to /56 or /48.
On IPv4 the opposite problem bites: many people share one address. Carrier-grade NAT puts a neighborhood behind one IP. The API flags it as threat.is_cgnat, and scoring caps the risk for CGNAT addresses. iCloud Private Relay does the same for Apple users: our relay IP scored 9, allow, with method relay_ip.
Relay geolocation needs care too. Our lookup placed 104.28.46.99 in Luanda, Angola. That's the relay's egress, not proof of where the user is, so don't geo-block on it. Our guide to detecting anonymized traffic without blocking real customers explains Apple's published egress list and how to treat relays.
Where to start with bot detection by IP
To act on bot detection by IP, map each signal to a default action: allow verified crawlers, rate-limit unverified crawler claims, block or step up on Tor, and use behavior signals where the IP looks residential. Each row below is one signal, what it proves, a default action and the post that covers it.
The fastest way to see where your own traffic sits: take the 20 busiest IPs from yesterday's logs and paste each into the free IP lookup tool, or run curl -s https://api.geoiphub.com/v1/lookup/<ip> without a key (30 lookups a minute), or 1,500 a day with a free key. Sort them by the bot classes in the first table, and you'll know which section of this map to read first.
Frequently Asked Questions
Can you detect bots by IP address alone?
Partly. An IP lookup tells you what kind of network a request came from: a published crawler range, a datacenter server, a Tor exit, a VPN, a privacy relay or a known scanner. It can't tell you whether a person or a script sent the request, and it can't spot a home IP rented out through a residential proxy network. Use the IP as one input next to request rate, account history and device signals.
How do I check whether a bot IP address is a real crawler?
Match the IP against the operator's published range file, or run a reverse DNS lookup and confirm it with a forward lookup. The User-Agent proves nothing because any client can send any string. In our lookup on 2026-10-07, 66.249.66.1 returned threat.crawler.verified true with name Googlebot, while a DigitalOcean server returned no crawler identity at all. Our crawler verification guide covers reverse DNS, range files and Web Bot Auth step by step.
How accurate is bot detection by IP?
It depends on the signal. A match against a published list is strong evidence: GeoIPHub checks 9 official crawler range feeds, refreshed every 2 days, and the Tor Project's exit list, refreshed daily. Hosting, VPN and proxy flags are inferences; hosting detection uses 890 curated hosting ASNs plus official ranges from 10 cloud providers. Residential proxies are not detected. We don't publish an accuracy percentage, and the five lookups in this post are examples, not a benchmark.
Does a high fraud score mean an IP is a bot?
No. The score measures risk, not automation. In our five lookups a Censys scanner scored 0 and a DigitalOcean server scored 15, both allow, although both are clearly automated. A Tor exit scored 100, block. Read the flags for what the traffic is, and the score for how risky it is.
Does GeoIPHub detect residential proxies?
No. The is_residential_proxy field is a coarse ASN-based inference, not a detection. A home IP rented out through a proxy network looks like an ordinary home connection at the IP layer, so catching it needs behavior and session signals.