One request from a scraper is just a GET. It has a status code, a path and a User-Agent, like every other line in your log. Scraping only shows up across hundreds of lines and across a network: a Python script on a cloud server, a headless browser driven by Puppeteer or Playwright, an AI training crawler, an SEO crawler like AhrefsBot, or a scraping API that sells rotating proxies by the gigabyte.
We ran seven scraper-class IPs through the GeoIPHub API on 2026-10-08. Four verified crawlers (Googlebot, CCBot, ClaudeBot, AhrefsBot) all came from hosting networks and scored allow. A plain Linode cloud server scored 10, also allow, and a Tor exit and a Mullvad VPN server both scored 100, block. The IP gives you the network. The scraping verdict comes from your logs.
Below we show how to detect scraping bots in an access log, trace one log line that claims to be Googlebot, show where IP data goes blind, and map what you find to allow, rate-limit, challenge or block. If you want the IP signals one by one first, our hub covers bot detection by IP, signal by signal; this post applies them to scraping.
How to detect web scraping: rate, path, client and network
To detect web scraping, check four groups of signals: how fast a client asks (rate), what it asks for and in what order (path), what it says about itself (client), and where it connects from (network). OWASP lists scraping as its own automated threat, OAT-011, defined as collecting "application content and/or other data for use elsewhere" (OWASP Automated Threats). That definition is about intent, which you can't see. The four groups are what you can see.
- Rate. Requests per minute per IP, and per network. By network we mean the ASN, or for finer grouping the /24 on IPv4 and the /64 on IPv6. Scrapers often run at a flat rate around the clock; people come in bursts and sleep.
- Path. Listing and pagination URLs walked in order, a sitemap read top to bottom, product or price pages only, and no CSS, JavaScript or image requests at all.
- Client. The User-Agent claim, a missing
Accept-Languageheader, no cookies kept between requests, and automation markers. In the browser,navigator.webdriveris true when Chrome runs with--headlessor--enable-automation(MDN), though a determined scraper patches it out. At the edge, a TLS fingerprint (JA3 or JA4) often shows an HTTP library behind a User-Agent that claims to be Chrome. - Network. What an IP lookup returns: a datacenter (hosting) IP, which the API flags as
detection.is_hosting, a verified crawler, Tor, a VPN or a scanner.
The hard part of web scraping detection is that every signal has an innocent twin. A CDN pulling your origin, an uptime monitor, an office behind one NAT address and an RSS reader all look odd on one axis. The case for scraping comes from signals that agree.
How to tell if your website is being scraped: read your access logs first
To tell if your website is being scraped, group your access logs by network and look for clients that fetch many pages, fetch them in order, and never load the assets a browser would. The outside signs tend to arrive first: your product copy or prices show up on another site, or bandwidth and server cost climb on listing pages while conversions stay flat. To check from outside, search for a distinctive sentence from your page in quotes, and watch for your prices reappearing on aggregator sites within hours of a change. The logs tell you who did it.
Here is what that looks like in a combined-format access log. These are illustrative values, not real traffic. 45.33.32.156 is Nmap's public test host on Linode, and it did not send these requests; it stands in for "a cloud server" because we may not publish a real visitor's IP.
# illustrative values, not real traffic
45.33.32.156 - - [08/Oct/2026:03:14:07 +0000] "GET /products?page=34 HTTP/1.1" 200 48211 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
45.33.32.156 - - [08/Oct/2026:03:14:08 +0000] "GET /products?page=35 HTTP/1.1" 200 47960 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
45.33.32.156 - - [08/Oct/2026:03:14:10 +0000] "GET /products?page=36 HTTP/1.1" 200 48377 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
45.33.32.156 - - [08/Oct/2026:03:14:11 +0000] "GET /products?page=37 HTTP/1.1" 200 48102 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
45.33.32.156 - - [08/Oct/2026:03:14:13 +0000] "GET /products?page=38 HTTP/1.1" 200 47815 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
45.33.32.156 - - [08/Oct/2026:03:14:14 +0000] "GET /products?page=39 HTTP/1.1" 200 48290 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
Three things stand out. The pages climb one by one, 1 to 2 seconds apart, at 3 a.m. There is no referrer and not one .css or .js request between them. And the client calls itself Googlebot, which the next sections test.
You don't need a log platform to find clients like this. These one-liners work on a standard combined log; run them over one busy hour first:
# Busiest client IPs
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -20
# Busiest IPv4 /24 networks (rotating scrapers often stay inside one block)
awk '{split($1,o,"."); print o[1]"."o[2]"."o[3]".0/24"}' access.log | sort | uniq -c | sort -rn | head -20
# Busiest User-Agents
awk -F'"' '{print $6}' access.log | sort | uniq -c | sort -rn | head -20
# Clients with more than 100 page requests and zero asset requests
awk '{ if ($7 ~ /\.(css|js|png|jpe?g|webp|svg|woff2?)(\?|$)/) a[$1]++; else p[$1]++ }
END { for (ip in p) if (!(ip in a) && p[ip] > 100) print p[ip], ip }' access.log | sort -rn
Work through the output with this list:
- Pages hit. Are the busy clients concentrated on listings, prices, search results or the sitemap?
- Rate per IP, then per network. After you enrich the IPs, add up requests per ASN; 50 hosting addresses in one ASN walking the same pages are likely one actor.
- Session depth. Hundreds of pages per session, with no cart, login or search box use.
- Order. Pagination or sitemap URLs fetched in sequence.
- Assets. HTML only, no CSS, JavaScript, fonts or images.
- Crawler claims. Every Googlebot, Bingbot, GPTBot or ClaudeBot User-Agent goes on a list to verify by IP.
- Automation markers. Library User-Agents (
python-requests,curl,Go-http-client) and headless browser hints.
Scraper IP addresses: what real GeoIPHub lookups return
Scraper IP addresses fall into a few network classes, and a lookup tells you which class you are looking at, not whether the request was a scraper. In our GeoIPHub lookups, verified crawlers scored 0 to 23 (allow); Tor and VPN exits scored 100 (block). We ran one anonymous lookup per class on 2026-10-08, four seconds apart:
curl -s https://api.geoiphub.com/v1/lookup/18.97.14.81
These are single lookups on one day, not a benchmark, and values move as ranges change. One known gap: no_ptr_datacenter on the Googlebot and AhrefsBot rows is our data, not theirs. The API stores no PTR record for some crawler ranges even though reverse DNS resolves (192.178.4.40 resolves to crawl-192-178-4-40.googlebot.com). For the ClaudeBot IP the signal is accurate: it has no reverse DNS.
Here is the CCBot response, trimmed but with real values and the original nesting:
{
"ip": "18.97.14.81",
"asn": {
"asn": 14618,
"org": "AMAZON-AES - Amazon.com, Inc.",
"asn_type": "hosting"
},
"detection": { "is_hosting": true, "is_vpn": false, "is_tor": false },
"threat": {
"is_crawler": true,
"crawler": {
"verified": true,
"name": "CCBot",
"operator": "CommonCrawl",
"category": "archiver",
"verification_method": "ip_range_list"
}
},
"scoring": {
"fraud_score": 0,
"recommended_action": "allow",
"detection_methods": ["datacenter_ip", "rdns_hosting", "verified_crawler"]
}
}
Common Crawl's own crawler runs on AWS, and still scores 0. The IP sits in 18.97.14.80/29, a range Common Crawl publishes in its ccbot.json file, and the verified-crawler adjustment in our scoring (−30) cancels the datacenter signal. GeoIPHub checks 19 official crawler range feeds this way, including Google's, Bing's, OpenAI's, Anthropic's, Common Crawl's and Ahrefs'. Whether you want a crawler whose data feeds AI models, like CCBot or GPTBot, reading your pages is a policy question. The lookup only settles who it is.
AhrefsBot was the surprise. It came back verified (5.39.1.230 sits in 5.39.1.224/27 in Ahrefs' published crawler ranges), but also is_abusive: true with abusive_listed in the methods, for a score of 23, still allow. We haven't traced which list flagged it. The lesson holds either way: "verified" tells you who a crawler is, not whether you want it. An SEO crawler that maps your site for a competitor's backlink report is still scraping, just declared.
Crawler impersonation in scraping logs: one line, traced
A request that says Googlebot in its User-Agent but comes from a cloud server is an unverified claim. GeoIPHub can tell you the IP isn't on a Google range, but the comparison with the User-Agent happens in your code, because the lookup never sees your headers. Take the fourth line of the illustrative log: 45.33.32.156, GET /products?page=37, User-Agent Googlebot/2.1.
- Look up the IP. The real lookup returned ASN 63949, Akamai Connected Cloud (Linode),
detection.is_hosting: true,threat.crawler.name: "none",verified: false,fraud_score10,allow, with methods["datacenter_ip"]. - Compare the claim. Your code checks the User-Agent token against
threat.crawler.verifiedandthreat.crawler.name. - Decide. The UA says Googlebot; the IP isn't on any Google range. That makes it "not verified as Googlebot". Rate-limit or challenge it. A challenge (a CAPTCHA or JavaScript check) is how you act on the API's
step_upaction. Don't call it fake: a new Google range can take up to about two days to verify in our data, because range files are fetched daily and the lookup database is rebuilt daily. - Check the contrast. Real Googlebot at
192.178.4.40returnedverified: true,name: "Googlebot",verification_method: "ip_range_list", score 0.
Google documents two ways to do the same check yourself: reverse DNS on the IP followed by a forward lookup of the hostname, or a match against its published range files such as common-crawlers.json (Google Search Central). The comparison in code looks like this:
// Map User-Agent tokens to the crawler names the API returns.
// They don't always match: OpenAI's "ChatGPT-User" comes back as "ChatGPTUser".
const CRAWLER_NAMES = {
"Googlebot": "Googlebot",
"GPTBot": "GPTBot",
"CCBot": "CCBot",
"ClaudeBot": "ClaudeBot",
"AhrefsBot": "AhrefsBot",
"ChatGPT-User": "ChatGPTUser",
};
async function classifyRequest(ip, userAgent) {
const res = await fetch(`https://api.geoiphub.com/v1/lookup/${ip}`, {
headers: { "x-api-key": process.env.GEOIPHUB_API_KEY },
});
const { threat, scoring } = await res.json();
const token = Object.keys(CRAWLER_NAMES).find((t) => userAgent.includes(t));
if (token) {
const ok = threat.crawler.verified && threat.crawler.name === CRAWLER_NAMES[token];
return ok ? "allow_crawler" : "unverified_crawler_claim"; // rate-limit or challenge
}
return scoring.recommended_action; // allow, review, step_up or block
}
This is the case where the score says allow and the pattern says scraper. On the IP alone, this request is 10 and allow, because a cloud server is a weak signal by itself. The scraping verdict comes from three things the score can't see: the UA mismatch, the steady rate and the page-by-page walk. For the full Google procedure, see how to verify Googlebot by reverse DNS and Google's range files; for GPTBot, ClaudeBot and signed agents, see how to verify AI crawlers by IP, reverse DNS and Web Bot Auth.
What IP data misses: residential proxies and real browsers
IP data misses scrapers that route through residential or mobile proxies, and GeoIPHub does not detect residential proxies; the is_residential_proxy field is a coarse ASN-based inference, so treat it as no signal. A residential proxy network sends each request through someone's home connection. At the IP layer the request looks like that household: a consumer ISP, a real city, no hosting flag, a clean score.
Rotation makes it worse. A pool can give every request a new home IP, so per-IP counters never climb. Mobile proxies share carrier IPs with real phones, and an IPv6 scraper can cycle through a whole /64 without touching a new network.
Real browsers are the second blind spot. A Playwright script driving Chrome from an office or home connection sends the same IP, the same TLS handshake and the same assets as the person at the next desk. Some headless setups announce themselves; tuned ones don't.
What still gives these scrapers away is the pattern across requests:
If you set a hidden link, disallow its path in robots.txt, so crawlers that honor the file never touch it. Two posts go deeper on the layers that sit beside the IP. Read why device fingerprinting and IP intelligence work best together, and what a lookup can and can't show in residential proxy fraud detection.
Scraping vs crawling vs API abuse
Crawling is declared and verifiable, scraping takes your content or data and is often undeclared, and API abuse is scripted calls to your endpoints beyond your terms or limits. The line between them is behavior and identity, not technology: Googlebot and a price scraper can run the same HTTP library on the same cloud.
A verified crawler can still scrape, as AhrefsBot shows, and a scraper can still look like a crawler. That's why the next section starts by allowing only the verified ones.
How to prevent web scraping without blocking search crawlers
To prevent web scraping without blocking search crawlers, allow verified crawlers first, then rate-limit by network and challenge clients that don't behave like browsers, and block only what you've confirmed. GeoIPHub's score bands give the IP side of that: allow up to 25, review up to 50, step_up up to 75 and block above 75. The order matters. A blanket datacenter block run first would have hit Googlebot, CCBot, ClaudeBot and AhrefsBot in our table, since all four came from hosting networks.
- Allow verified crawlers when
threat.crawler.verifiedis true andcrawler.namematches the User-Agent token. Then apply your own policy per crawler. For AI crawlers, our comparison of GPTBot, OAI-SearchBot and ChatGPT-User shows how to block training without leaving AI search. - Reuse the IP rules. Score bands, Tor handling and per-prefix rate limits work the same for scrapers as for any bot; the hub covers them. Treat a datacenter IP as a weighted input, not a block, since hosting also covers AI agents, corporate proxies and your own tools (see what a datacenter IP address is).
- Rate-limit by ASN on the pages scrapers want. A pool rotating through one hosting provider stays under any per-IP limit, but a per-ASN counter on listing, search and price pages adds it up. Give large clouds a higher ceiling, since many unrelated customers share one ASN.
- Cap pagination and put full listings behind a login. Serve a fixed number of listing pages to anonymous visitors. An account wall turns each scraper into accounts you can count and close.
- Set per-key quotas on the JSON endpoints your pages call. Scrapers often skip the HTML and call the API behind it, so issue keys or session tokens and cap calls per key.
- Add a bot management or WAF layer for the client signals your logs don't hold, such as TLS fingerprints and JavaScript checks at the edge.
- Use behavior rules for residential traffic, where the IP gives you nothing.
The responses form a ladder: allow, rate-limit, challenge, block. Some sites add a tarpit, which slows answers to a suspected scraper, or serve decoy content, such as altered prices, to traffic they have already classified. Both cost you something if the classification is wrong, so use them only on traffic you'd block anyway.
For the IP-only rows (Tor, relays, scanners, IPv6 prefixes), use the action map in our bot detection by IP guide.
And robots.txt? It asks; it doesn't enforce. RFC 9309, the robots exclusion standard, says its rules "are not a form of access authorization" (RFC 9309). Keep it for the crawlers that honor it, and enforce everything else on your server.
Start with yesterday's logs. Export the 50 busiest clients that loaded no assets, paste each into the free IP lookup (30 lookups a minute without a key, 1,500 a day with a free one), and sort them into the classes in the lookup table above. The ones left over, busy and on consumer networks, are where your behavior rules go next.
Frequently Asked Questions
How do I block AI scrapers like GPTBot or CCBot?
Start with robots.txt: GPTBot and CCBot both honor it, so a Disallow rule for their tokens stops the declared crawlers. robots.txt is a request, not access control (RFC 9309 says its rules are not a form of access authorization), so enforce on your server too: verify each crawler by its published IP ranges and treat a client that uses the name from outside those ranges as unverified. OpenAI runs several bots with different jobs; our post GPTBot vs OAI-SearchBot vs ChatGPT-User explains which one to block if you want to stay in ChatGPT search.
How many requests per minute means scraping?
There is no universal number. Measure your own human baseline on a normal day, per /24 on IPv4 and per /64 on IPv6, then look at networks far above it. Rate alone also flags office NAT, CGNAT and monitoring services, so treat a high rate as scraping only when it comes with pages fetched in order and no asset requests.
Can you stop web scraping completely?
No. Anything a browser can render, a determined scraper can copy, including through residential proxies and real browsers. The goal is to make scraping slow, expensive and visible: verify crawlers, rate-limit by network, challenge clients that don't behave like browsers, and put high-value data behind a login or API keys.
Does blocking datacenter IPs stop scraping?
Partly. It stops scrapers running on cheap cloud servers, but it also blocks verified crawlers unless you allow them first: in our lookups on 2026-10-08, CCBot on AWS scored 0 (allow) and real Googlebot scored 0 (allow), both on hosting networks. A plain Linode server scored 10, also allow, so the score alone won't block it. Scrapers that rent residential proxies get past a datacenter block entirely.
Can an IP lookup tell me an IP is a scraper?
No. A lookup tells you the network class and the risk: a verified crawler, a datacenter server, a Tor exit or a VPN. Whether that IP is scraping you comes from your logs: its request rate, the order of the pages it fetches, whether it loads assets, and whether its User-Agent matches the crawler the lookup verifies.
How do I know if a scraper is using residential proxies?
You can't tell from the IP, and GeoIPHub does not detect residential proxies. Look for many consumer IPs that share one session pattern: the same page order, the same timing between requests, the same missing assets and similar TLS or device fingerprints. The pattern across addresses gives the scraper away, not any single address.