Skip to content
Network Security

How to Verify Googlebot by IP and Spot Fake Googlebots

8 min readGeoIPHub Team

The User-Agent is the part an impostor controls. In a 14-day test on one new site published by Search Engine Journal, only 107 of 799 requests calling themselves Googlebot came from a verified Google address. That's one small site, so read it as a warning, not a rate. Below is the manual check, the automated one, and a real lookup of a true and a fake Googlebot side by side.

Why a Googlebot User-Agent proves nothing

A User-Agent string is a claim the client makes about itself, and any client can send any string. Scrapers copy Googlebot's because many sites treat it kindly: no rate limits, no login walls, no bot challenges. That makes the name a disguise worth wearing.

The cost of believing it runs both ways. Let a fake Googlebot through and a scraper gets your content at crawler speed. Block a real one by mistake and your pages drop out of Google's index. The fix is the same for both: verify the address. Our broader guide to verifying AI crawlers by IP covers GPTBot, ClaudeBot and signed agents, and our comparison of OpenAI's main bots, GPTBot, OAI-SearchBot and ChatGPT-User covers which of them to block; this post covers Google's crawlers only.

How to verify Googlebot manually: reverse DNS, then forward DNS

Google documents a two-step DNS check, and both steps matter. Here it is on a real Googlebot address, run on 2026-10-05:

GeoIPHub API
# 1. Reverse lookup: which hostname does the IP claim? dig +short -x 66.249.66.1 # crawl-66-249-66-1.googlebot.com. # 2. Forward lookup: does that hostname point back to the same IP? dig +short crawl-66-249-66-1.googlebot.com # 66.249.66.1

The hostname ends in googlebot.com and resolves back to the original IP, so the request came from Google. Skipping step 2 is the common mistake. Whoever controls an IP also controls its reverse DNS record, so a scraper can name its own server crawl-fake.googlebot.com. It can't make Google's DNS point that name back at its server. The forward lookup is what makes the check unforgeable.

This works well for one address you're investigating. It's slow for a busy log, because every new IP needs two DNS queries, and DNS failures turn into false "fakes".

Googlebot IP ranges: three crawler types, three range files

Google's crawlers fall into three groups, and each has its own hostnames and its own published IP list. Matching an IP against the right file is the automated version of the DNS check, and it's how most verification at scale is done.

Two details trip up older scripts. First, the files moved: the old googlebot.json URL now answers with a 301 redirect to /static/crawling/ipranges/common-crawlers.json, so code that doesn't follow redirects silently gets nothing. Second, Google now refreshes these files daily. The copy we fetched on 2026-10-05 was generated on 2026-10-02 and lists 317 prefixes, 170 of them IPv4. Cache the file, but refresh it at least daily.

Source for the table: Google's Verify requests from Google crawlers and fetchers page.

A real Googlebot and a fake one, side by side

The quickest way to see the difference is to look up both addresses. We ran two lookups on 2026-10-05 (results change as ranges move). The first is 66.249.66.1, the Googlebot address from Google's own documentation. The second is 104.248.50.21, a DigitalOcean server: imagine your log shows it sending a Googlebot User-Agent.

GeoIPHub API
[ { "ip": "66.249.66.1", "asn": { "asn": 15169, "org": "GOOGLE - Google LLC", "asn_type": "hosting", "user_type": "search_engine_spider" }, "detection": { "is_hosting": true, "is_vpn": false, "is_proxy": false, "is_tor": false }, "threat": { "is_crawler": true, "crawler": { "verified": true, "name": "Googlebot", "operator": "Google", "category": "search_engine", "verification_method": "ip_range_list" } }, "scoring": { "fraud_score": 0, "recommended_action": "allow", "detection_methods": ["datacenter_ip", "no_ptr_datacenter", "verified_crawler"] } }, { "ip": "104.248.50.21", "asn": { "asn": 14061, "org": "DIGITALOCEAN-ASN - DigitalOcean, LLC", "asn_type": "hosting", "user_type": "hosting" }, "detection": { "is_hosting": true, "is_vpn": false, "is_proxy": false, "is_tor": false }, "threat": { "is_crawler": false, "crawler": { "verified": false, "name": "none", "verification_method": "none" } }, "scoring": { "fraud_score": 15, "recommended_action": "allow", "detection_methods": ["datacenter_ip", "no_ptr_datacenter"] } } ]

The first address is inside Google's range list, so threat.crawler.verified is true and verification_method says how: ip_range_list. The verified_crawler signal pulls its fraud_score to 0. The second is an ordinary cloud server. It isn't a crawler of any kind, so a request from it carrying a Googlebot User-Agent is a fake by definition.

Notice what the score does not do. The fake scores 15, allow, because a cloud server on its own isn't malicious. The verdict "fake Googlebot" comes from combining two facts: your log says Googlebot, the lookup says not Google. The API sees the IP, not your headers, so that comparison is yours to make:

GeoIPHub API
const claimsGooglebot = /Googlebot/i.test(req.headers["user-agent"] ?? ""); const res = await fetch(`https://api.geoiphub.com/v1/lookup/${clientIp}`, { headers: { "x-api-key": process.env.GEOIPHUB_KEY }, }); const ip = await res.json(); const isRealGoogle = ip.threat.crawler.verified && ip.threat.crawler.operator === "Google"; if (claimsGooglebot && !isRealGoogle) { // Treat as an ordinary unverified bot: rate-limit or challenge it. }

You can try both addresses in the free IP lookup tool without a key.

What to do with a fake Googlebot

Treat a failed check as "unverified bot", not as an attack. The request is from someone pretending to be Google, which tells you they want crawler treatment. Remove that treatment and handle them like any other automated client.

  • Drop the crawler allowances. No bypass of rate limits, paywalls or bot challenges for a Googlebot name that didn't verify.
  • Rate-limit by network, not by address. Scrapers rotate IPs inside one provider. For IPv6 sources, key on the prefix, as in our guide to rate-limiting IPv6 by /64 prefix.
  • Never block a verified Google address by accident. Run the verification before any blocklist rule, so a broad hosting or ASN block can't catch the real Googlebot.
  • Verify the real client IP. Use the connecting address, or a forwarding header only when your own proxy or CDN sets it. A client can write anything into X-Forwarded-For, including a real Googlebot address.
  • Fail open on DNS errors. If the reverse lookup times out, that's not proof of a fake. Fall back to the range file, or let the request through and check later.
  • Keep the evidence. Log the claimed User-Agent, the IP and the verification result, so you can measure how much "Googlebot" traffic is real.

The same pattern applies at checkout and signup, where automated agents increasingly claim to be well-known bots. Our guide to blocking AI shopping-agent fraud at checkout shows how a verified or unverified crawler flag fits into a payment decision.

Where IP verification falls short

Range matching tells you an address belongs to Google's crawlers. It has limits worth knowing before you rely on it.

  • Our coverage is Google's common and special-case crawlers only. GeoIPHub verifies crawlers from 9 official range feeds: Googlebot, Google Special Crawlers, Bingbot, Applebot, DuckDuckBot, GPTBot, ChatGPT-User, OAI-SearchBot and CCBot. Google's user-triggered fetchers aren't in that list, so for those, use Google's own files or the DNS check.
  • There can be a short lag. Google refreshes its files daily and we refresh crawler feeds every 2 days, so a brand-new Google prefix can take up to a couple of days to verify through us. The DNS check has no lag.
  • Our record holds no reverse DNS for these addresses. That's why no_ptr_datacenter appears on the real Googlebot above, even though dig returns crawl-66-249-66-1.googlebot.com. The verdict comes from the range list, which is the reliable path, but the method label is misleading. We've logged it as a bug.
  • Signed requests are coming, but not for Googlebot yet. Google is testing Web Bot Auth, which lets an agent cryptographically sign its requests, with some AI agents on its infrastructure. It doesn't sign every request, and Google says to fall back to the established checks above. Our AI crawler verification guide covers how to verify a signature.
  • A real Google address isn't always Googlebot. Google's network, AS15169, also carries services such as Google Public DNS (8.8.8.8), which our lookup scores as plain hosting, not a crawler. Only the crawler range files count, so never allowlist AS15169 as a whole.

Start with one number from your own logs: how many requests last week claimed to be Googlebot, and how many of those IPs pass the forward-confirmed DNS check. The gap between the two is your fake-Googlebot traffic, and it's usually bigger than people expect.

Frequently Asked Questions

What is Googlebot's IP address?

Googlebot has no single IP address. It crawls from many prefixes that Google publishes in common-crawlers.json, refreshed daily; the copy we fetched on 2026-10-05 listed 317 prefixes, 170 of them IPv4. 66.249.66.1, the example in Google's own documentation, is one of them. Always check an address against the current file instead of a hard-coded list.

How do I verify that a request is really from Googlebot?

Check the IP, not the User-Agent. Run a reverse DNS lookup on the IP and confirm the hostname ends in googlebot.com, google.com or googleusercontent.com, then run a forward lookup on that hostname and confirm it returns the same IP. At scale, match the IP against Google's published range files instead, starting with common-crawlers.json for Googlebot.

Where does Google publish Googlebot's IP ranges?

In JSON files linked from Google's 'Verify requests from Google crawlers and fetchers' page. Googlebot and the other common crawlers are in common-crawlers.json. The old googlebot.json address now returns a 301 redirect to that file. Special-case crawlers such as AdsBot and user-triggered fetchers have their own files.

Is every Google IP a Googlebot IP?

No. Google's network (AS15169) also carries other Google services, such as Google Public DNS at 8.8.8.8. Only addresses inside the crawler range files are Google crawlers, so an ASN allowlist is too broad. In our lookup, 66.249.66.1 verified as Googlebot while a DigitalOcean server claiming the same name did not.

Should I block requests that claim to be Googlebot but fail verification?

Treat them as an ordinary unverified client, not as Googlebot. Rate-limit or challenge them like any other bot. Don't give them the crawl allowances you give the real Googlebot, and check that your own rules can never block a verified Google address.