Skip to content
Network Security

GPTBot vs OAI-SearchBot vs ChatGPT-User: Which to Block

16 min readGeoIPHub Team

On 2026-10-06, OpenAI's range files listed 18 prefixes for GPTBot, 39 for OAI-SearchBot and 230 for ChatGPT-User. OpenAI documents a fourth bot, OAI-AdsBot, but it only visits pages submitted as ChatGPT ads, so the three above are the ones most sites have to decide on. A rule that treats them as one thing called "OpenAI" either gives away training rights you meant to keep or drops you out of ChatGPT search by accident.

Below is what each bot does in OpenAI's own words, what blocking each one costs, the robots.txt lines for the most common policy, and how to check that a request claiming to be one of them came from OpenAI. We looked up one IP from each range file, and one impostor, so you can see the difference in real API output. The general method behind that check is in our guide to verifying AI crawlers by IP, reverse DNS and Web Bot Auth.

GPTBot vs OAI-SearchBot vs ChatGPT-User: what each one does

GPTBot is for training, OAI-SearchBot is for search, and ChatGPT-User is for a person's live request. That's the whole split, and OpenAI's crawler documentation states it directly: "Each setting is independent of the others." Here are the three side by side, quoted from that page and from the range files we fetched on 2026-10-06.

These are the full User-Agent strings as OpenAI's page showed them on 2026-10-06:

GeoIPHub API
GPTBot: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot OAI-SearchBot: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot ChatGPT-User: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

OpenAI labels the GPTBot and OAI-SearchBot strings as examples and notes "the version number may change"; the ChatGPT-User string is listed as the full string. OAI-SearchBot's string begins like desktop Chrome on a Mac, so match on the token (GPTBot, OAI-SearchBot, ChatGPT-User), not the whole string. One detail helps with log analysis: when GPTBot or OAI-SearchBot fetch your robots.txt, OpenAI says they may add a robots.txt marker to the User-Agent, so you can tell those requests apart even when your logs don't record paths.

OpenAI's page also lists a fourth bot, OAI-AdsBot. It is "used to validate the safety of web pages submitted as ads on ChatGPT", and OpenAI says it "may also use content from the landing page to determine when it's most relevant to show the ad to users." It only visits pages submitted as ads, and OpenAI says the data it collects is not used to train its foundation models. Its token is OAI-AdsBot/1.0, and its range file, openai.com/adsbot.json, was created on 2026-05-12 and listed 2 IPv4 prefixes when we fetched it. OpenAI doesn't list OAI-AdsBot as a robots.txt tag: the page names only OAI-SearchBot and GPTBot as those.

ChatGPT-User vs GPTBot: why a user-triggered fetch is different

ChatGPT-User differs from GPTBot because a person, not a crawl schedule, decides when it visits. OpenAI's page says: "ChatGPT-User is not used for crawling the web in an automatic fashion. Because these actions are initiated by a user, robots.txt rules may not apply."

Two more sentences from the same entry matter for policy. ChatGPT-User "is not used to determine whether content may appear in Search", and OpenAI points site owners to OAI-SearchBot instead: "Please use OAI-SearchBot in robots.txt for managing Search opt outs and automatic crawl."

So a Disallow line for ChatGPT-User is, at best, a request OpenAI says it may not honour. If you truly need to stop user-triggered fetches, that has to happen at your firewall or application, by IP. We looked for a dated changelog of when this wording was introduced and didn't find one: the page links an RSS feed for updates, and it was empty when we checked. So we quote the current text and make no claim about when it changed.

ChatGPT-User isn't the only way ChatGPT reaches a site. OpenAI's help article on allowlisting ChatGPT Work's Cloud browser (the URL still says "chatgpt-agent", and Akamai still files it under the bot name "ChatGPT Agent") says the Cloud browser "uses Web Bot Auth to sign outbound HTTP requests". Each request carries Signature and Signature-Input headers plus a Signature-Agent header set to "https://chatgpt.com", and you verify it under RFC 9421 against the keys at chatgpt.com/.well-known/http-message-signatures-directory. That traffic is identified by its signature, not by a range file. Our pillar guide's Web Bot Auth section covers how signature checks work.

GPTBot, OAI-SearchBot and ChatGPT-User IP ranges

OpenAI publishes the three bots' IP ranges at https://openai.com/gptbot.json, https://openai.com/searchbot.json and https://openai.com/chatgpt-user.json. As of 2026-10-06 they listed 18, 39 and 230 IPv4 prefixes, and no IPv6. Each range file is a small JSON document with a creationTime and a list of prefixes. Here is the start of the GPTBot file, and a one-line way to summarise all three:

GeoIPHub API
curl -s https://openai.com/gptbot.json | head -c 120 # { # "creationTime": "2026-09-22T02:00:07.000000", # "prefixes": [ # { # "ipv4Prefix": "132.196.86.0/24" for f in gptbot searchbot chatgpt-user; do curl -s https://openai.com/$f.json | jq -c '[.creationTime, (.prefixes | length), .prefixes[0]]' done # ["2026-09-22T02:00:07.000000",18,{"ipv4Prefix":"132.196.86.0/24"}] # ["2026-01-02T11:00:00.000000",39,{"ipv4Prefix":"104.210.140.128/28"}] # ["2026-09-25T18:04:57.257335",230,{"ipv4Prefix":"104.208.184.192/28"}]

Three things stand out on the day we fetched them:

  • All three range files are IPv4 only. None contained an ipv6Prefix entry. Code that only parses ipv4Prefix works today, but read both keys, since the shape allows IPv6.
  • The shapes are very different. GPTBot is 18 prefixes, mostly /25. ChatGPT-User is 229 tiny /28 blocks of 16 addresses plus one large /17, 9.129.0.0/17. A hand-kept ChatGPT-User allowlist has the most entries to get wrong.
  • All three OpenAI IPs we sampled (one per file) sit on Microsoft's network (AS8075). That's exactly why you can't allowlist the ASN: the same network carries every other Azure customer, including anyone renting a server to impersonate GPTBot.

Copied lists rot. OpenAI updates these range files, and two of the three carry a creationTime from the two weeks before we fetched them, so a list pasted from a blog post (including this one) is out of date the next time a file changes. Fetch the file, cache it, and refresh it on a schedule. Don't paste the ranges into a config file. Prefix matching itself, including how to match an IP against published CIDR ranges per bot, is covered in the pillar guide; this post sticks to OpenAI's three bots.

Should I block GPTBot? What blocking each bot costs

Block GPTBot if you don't want your content used for training; it costs you nothing in ChatGPT search, according to OpenAI. Blocking the other two has a visible cost, so it should be a deliberate choice, not a side effect of a blanket rule.

One more line from OpenAI is worth knowing if you allow both crawlers: "If your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling." Allowing both doesn't necessarily double the load.

For context on what other sites do: an Ahrefs study of about 140 million websites' robots.txt files, published 2025-05-21 with data through December 2024, found GPTBot was the most blocked AI bot at 5.89% of sites. OAI-SearchBot was blocked by 5.54% and ChatGPT-User by 5.64%. Most of that is blanket rules: only 0.5% of sites named GPTBot explicitly, 0.22% ChatGPT-User and 0.08% OAI-SearchBot. Read these as context, not current rates. The data predates today's bot descriptions, and it covers robots.txt only, not firewall or IP blocks.

The gap between 5.54% and 0.08% is the point. Almost none of the sites blocking OAI-SearchBot named it. The block came from a general rule, such as User-agent: * with Disallow: /, or an allowlist that permits only certain bots, and either one catches the search bot along with the training bot.

To block GPTBot but allow ChatGPT search, disallow GPTBot and allow OAI-SearchBot in robots.txt, using the tokens exactly as OpenAI writes them:

GeoIPHub API
User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: / User-agent: * Allow: /

There's no ChatGPT-User group, because OpenAI says robots.txt "may not apply" to it. OpenAI also says search results can take about 24 hours to reflect a robots.txt change.

robots.txt is only half of what OpenAI asks for. Under OAI-SearchBot, its page recommends "allowing OAI-SearchBot in your site's robots.txt file and allowing requests from our published IP ranges". The IP half is easy to break without noticing. The OAI-SearchBot address we sampled is a Microsoft Azure IP that our API flags as datacenter_ip, so a WAF rule that blocks or challenges datacenter or Azure traffic can keep OAI-SearchBot out and drop you from ChatGPT search even when robots.txt says Allow. If you run a rule like that, exempt the prefixes in searchbot.json from it.

Check the file your server actually sends, not the one in your repository. A CDN can change it on the way out. Cloudflare's managed robots.txt, for example, "will prepend our managed robots.txt before your existing robots.txt". Its rules disallow GPTBot but don't mention OAI-SearchBot or ChatGPT-User, so on its own it produces the same "block training, allow search" policy as the file above. Any setting like this adds rules you didn't write, so test the served file:

GeoIPHub API
import urllib.robotparser p = urllib.robotparser.RobotFileParser("https://your-site.com/robots.txt") p.read() for bot in ["GPTBot", "OAI-SearchBot", "ChatGPT-User"]: print(bot, p.can_fetch(bot, "/blog/"))

Run against the file above, Python 3.9's parser printed GPTBot False, OAI-SearchBot True and ChatGPT-User True. If OAI-SearchBot comes back False, something upstream is blocking ChatGPT search.

robots.txt vs IP enforcement: a robots.txt rule is a request, not a block

A robots.txt rule tells honest crawlers what you'd like, but only an IP check where the request arrives actually stops a bot. OpenAI's real bots read robots.txt. A scraper that puts GPTBot or ChatGPT-User in its User-Agent ignores it, and borrowing the name costs one flag: curl -A "ChatGPT-User/1.0" https://your-site.com/ carries the same bot token as the real thing.

The User-Agent can't tell you which is which, and neither can robots.txt. If your firewall or CDN gives OpenAI's bots a pass on rate limits or bot challenges, that pass is only as good as your check of who is asking.

Enforcement has to happen where the request arrives, and it needs two facts together:

  1. The claim: which OpenAI token the User-Agent names.
  2. The proof: whether the connecting IP is inside that same bot's range file.

A claim with no proof is not verified as OpenAI: challenge or rate-limit it rather than treating it as an attack, because the pillar post explains why an IP outside a published range is inconclusive rather than proof of a fake. Pin each token to its own range file, because a ChatGPT-User IP doesn't prove a GPTBot claim. Reverse DNS doesn't help here: we ran dig -x on one IP from each file and all three returned NXDOMAIN from Microsoft's Azure DNS, so there's no hostname to forward-confirm. The range file is the check that works for OpenAI.

How to check a request really came from OpenAI: three real IPs and one fake

A request came from OpenAI when its IP is inside the range file for the bot its User-Agent names. Our API returns verified: true and that bot's name for an IP inside the file, and verified: false for an IP outside every crawler feed we check. We looked up the first address in the first prefix of each OpenAI range file on 2026-10-06, plus a DigitalOcean server. Results change as ranges move.

Note that the API's crawler names drop the hyphens: OAISearchBot and ChatGPTUser, not OAI-SearchBot and ChatGPT-User.

All three OpenAI addresses came back with operator: "OpenAI", category: "ai" and verification_method: "ip_range_list". Now the traced example. Imagine a log line, illustrative and not from our traffic: a request with a ChatGPT-User/1.0 User-Agent from 104.248.50.21. Here is that IP next to a real ChatGPT-User address, trimmed from the live responses:

GeoIPHub API
[ { "ip": "104.208.184.192", "asn": { "asn": 8075, "org": "MICROSOFT-CORP-MSN-AS-BLOCK - Microsoft Corporation", "asn_type": "hosting" }, "detection": { "is_hosting": true, "is_vpn": false, "is_proxy": false }, "threat": { "is_crawler": true, "crawler": { "verified": true, "spoofed": false, "name": "ChatGPTUser", "operator": "OpenAI", "category": "ai", "verification_method": "ip_range_list" } }, "scoring": { "fraud_score": 0, "recommended_action": "allow", "detection_methods": ["datacenter_ip", "no_ptr_datacenter", "verified_crawler"] } }, { "ip": "104.248.50.21", "asn": { "asn": 14061, "org": "DIGITALOCEAN-ASN - DigitalOcean, LLC", "asn_type": "hosting" }, "detection": { "is_hosting": true, "is_vpn": false, "is_proxy": false }, "threat": { "is_crawler": false, "crawler": { "verified": false, "spoofed": false, "name": "none", "operator": "none", "category": "none", "verification_method": "none" } }, "scoring": { "fraud_score": 15, "recommended_action": "allow", "detection_methods": ["datacenter_ip", "no_ptr_datacenter"] } } ]

Read it in three steps:

  1. The lookup. The DigitalOcean IP isn't in any crawler feed we check, so threat.crawler.verified is false and name is "none". Its score is 15, allow, because a datacenter IP on its own isn't malicious.
  2. The comparison. The API receives only the IP, so it never saw the ChatGPT-User claim, and spoofed: false here doesn't mean the request is honest. Your code holds both facts: the request claims an OpenAI token, and the verified identity is "none". That mismatch is the spoof.
  3. The contrast. 104.208.184.192 is in chatgpt-user.json, so it comes back verified: true, name: "ChatGPTUser", and verified_crawler pulls its score to 0.

The comparison is a few lines. The map translates each User-Agent token to the API's hyphen-free name:

GeoIPHub API
const OPENAI_BOTS = { "GPTBot": "GPTBot", "OAI-SearchBot": "OAISearchBot", "ChatGPT-User": "ChatGPTUser" }; const ua = req.headers["user-agent"] ?? ""; const claimed = Object.keys(OPENAI_BOTS).find((token) => ua.includes(token)); if (claimed) { const res = await fetch(`https://api.geoiphub.com/v1/lookup/${clientIp}`, { headers: { "x-api-key": process.env.GEOIPHUB_KEY }, }); const { threat } = await res.json(); const isThatBot = threat.crawler.verified && threat.crawler.name === OPENAI_BOTS[claimed]; if (!isThatBot) { // Claims an OpenAI bot, but the IP isn't in that bot's file: challenge or rate-limit it. } }

To try it without a key, run curl https://api.geoiphub.com/v1/lookup/104.208.184.192 (30 anonymous lookups a minute), or paste an IP from your own logs into the free IP lookup tool. For IPv6 sources in the "unverified" bucket, rate-limit by prefix, as in our guide to rate-limiting and blocking IPv6 by /64 prefix.

Where this check falls short

The check has three limits: it doesn't cover OAI-AdsBot or ChatGPT's signed Cloud browser traffic, it can't see your request headers, and a new prefix can take up to 2 days to verify through us. One output field is also wrong on our impostor sample.

  • We verify three of OpenAI's four bots. GeoIPHub's crawler verification checks 9 official range feeds: Googlebot, Google Special Crawlers, Bingbot, Applebot, DuckDuckBot, GPTBot, ChatGPT-User, OAI-SearchBot and CCBot. OAI-AdsBot isn't one of them, and neither are ClaudeBot or PerplexityBot. For those, use the operator's own range file.
  • We don't verify signatures. ChatGPT Work's Cloud browser is identified by its Web Bot Auth signature, which none of our 9 feeds cover. Verify it yourself against OpenAI's key directory; OpenAI's article says Cloudflare, Akamai and HUMAN verify it automatically.
  • The API sees an IP, not your headers. It can tell you an address belongs to ChatGPT-User. Matching that to the User-Agent the request claimed is your code's job, as above.
  • Feeds can lag by up to 2 days. We refresh crawler feeds every 2 days, and two of OpenAI's three range files were created in the fortnight before our fetch. A brand-new prefix may not verify through us for a couple of days, so a mismatch is a reason to challenge, not to hard-block.
  • no_ptr_datacenter is accurate here, and wrong on our impostor. It fired on all three OpenAI IPs, and dig -x confirms they have no reverse DNS. But 104.248.50.21 does have a live PTR record, which our record doesn't hold. That's the same gap we found when verifying Googlebot by IP and spotting fakes, and it's logged as a backend bug. The crawler verdict doesn't depend on it.

The same claim-versus-proof check matters at checkout, where an automated buyer can claim any bot name it likes. See how to handle AI shopping agents at checkout by IP for that version.

Start with your served robots.txt: run the short Python check above and make sure OAI-SearchBot says True unless you meant otherwise. Then pull a week of requests naming GPTBot, OAI-SearchBot or ChatGPT-User from your logs and check their IPs against the three range files. The ones that don't match are not verified as OpenAI: challenge or rate-limit them.

Frequently Asked Questions

Is ChatGPT-User the same as GPTBot?

No. GPTBot crawls on its own schedule and collects content that may be used to train OpenAI's models. ChatGPT-User visits a page only when a person using ChatGPT or a Custom GPT triggers it, and OpenAI says it is not used to crawl the web automatically. They also publish separate IP range files: gptbot.json and chatgpt-user.json.

Can I block ChatGPT-User with robots.txt?

Not reliably. OpenAI says ChatGPT-User fetches are initiated by a user, so "robots.txt rules may not apply." It also says ChatGPT-User is not used to decide whether content appears in ChatGPT search; OAI-SearchBot is. If you need to stop user-triggered fetches, block them by IP at your firewall or application, using the ranges in openai.com/chatgpt-user.json.

What is OAI-SearchBot?

OAI-SearchBot is OpenAI's search crawler. OpenAI says it is "used to surface websites in search results in ChatGPT's search features." Its User-Agent carries the token OAI-SearchBot/1.4 inside a string that looks like desktop Chrome, and its IP ranges are at openai.com/searchbot.json, which listed 39 IPv4 prefixes on 2026-10-06. OpenAI recommends allowing it in robots.txt and allowing requests from its published IP ranges.

Does OAI-SearchBot respect robots.txt?

Yes. OpenAI names OAI-SearchBot and GPTBot as its robots.txt tags and says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.

How long does a robots.txt change take for ChatGPT search?

About 24 hours, according to OpenAI, which says "it can take ~24 hours from a site's robots.txt update for our systems to adjust."

Where are GPTBot's IP ranges?

At https://openai.com/gptbot.json. The copy we fetched on 2026-10-06 was created on 2026-09-22 and listed 18 IPv4 prefixes and no IPv6. OAI-SearchBot's ranges are at openai.com/searchbot.json and ChatGPT-User's at openai.com/chatgpt-user.json. Fetch the files on a schedule instead of copying the ranges, because they change.

Why does GPTBot come from a Microsoft IP?

In our sample it did: the GPTBot address we looked up, and the OAI-SearchBot and ChatGPT-User addresses, all sat on Microsoft's network, AS8075. OpenAI's crawler page doesn't say which network it crawls from, so we only report what our sample showed. Because other Microsoft cloud customers share that ASN, it proves nothing on its own. Check the IP against gptbot.json instead.

Will blocking GPTBot remove me from ChatGPT?

Not from ChatGPT search, according to OpenAI. Its documentation says each setting is independent, and gives the example of allowing OAI-SearchBot to appear in search results while disallowing GPTBot to opt out of training. Blocking OAI-SearchBot is what removes you from ChatGPT search answers.