AI Crawlers: Which to Block, Which to Keep
GPTBot and OAI-SearchBot come from the same company and do opposite jobs for you. Here is every major AI crawler, its exact robots.txt token, and what blocking it costs.
Three kinds of AI crawler with three different trades. Decide per column, not per brand name.
"Block the AI bots" sounds like one decision. It is at least three.
Block the wrong token and you stay out of model training, which you may want, but you also vanish from ChatGPT search answers, which you almost certainly don't. This post sits under our AI search guide and deals only with the crawlers themselves: our robots.txt guide covers the file's syntax in general, and our llms.txt explainer covers that separate file.
We'll use one running example. Kesar Lane is a hypothetical Pune D2C brand selling ayurvedic skincare. It spent two years writing ingredient guides. The founder wants two things that sound contradictory: keep the guides out of the next model's training set, and still get recommended when someone asks Perplexity for "kumkumadi oil for dry skin".
Both are possible. Here's how.
What are the three kinds of AI crawler?
The split that matters is what the crawler does with your page after it fetches it. Vendors now document this themselves, and they document it separately for each bot.
Training crawlers collect pages that may go into future models. OpenAI describes GPTBot as the crawler used to "crawl content that may be used in training our generative AI foundation models" (OpenAI crawler docs). Blocking one of these doesn't pull you out of any live answer product. The model just doesn't learn from your future pages.
Search crawlers build the index an AI assistant searches when it answers a live question. These are the ones that put a link to your page next to the answer. Perplexity says PerplexityBot is "designed to surface and link websites in search results on Perplexity" and "is not used to crawl content for AI foundation models" (Perplexity bot docs).
User-triggered fetchers visit a page because a person asked the assistant to read it, or the assistant decided mid-answer that it needed to. They behave more like a browser than a crawler, and that changes how much robots.txt controls them. More on that below.
Every major AI crawler, and what blocking it costs you
This is the table to bookmark. The "token" column is the exact string you put after User-agent:. Matching is case-insensitive under RFC 9309, but copy the vendor's spelling anyway so your file stays readable.
| Crawler | Company | Purpose | robots.txt token | What blocking it costs you |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | GPTBot | Nothing visible. Future content is marked as not for training. |
| OAI-SearchBot | OpenAI | Search / retrieval | OAI-SearchBot | You won't be shown in ChatGPT search answers; navigational links can still appear. |
| ChatGPT-User | OpenAI | User-triggered | ChatGPT-User | OpenAI says robots.txt rules "may not apply", so blocking may do little. |
| OAI-AdsBot | OpenAI | Ad-page safety checks | Not controlled by robots.txt | Not applicable. It only checks pages submitted as ChatGPT ads. |
| ClaudeBot | Anthropic | Training | ClaudeBot | Future content excluded from Anthropic's training datasets. |
| Claude-SearchBot | Anthropic | Search / retrieval | Claude-SearchBot | "May reduce your site's visibility and accuracy in user search results." |
| Claude-User | Anthropic | User-triggered | Claude-User | "May reduce your site's visibility for user-directed web search." |
| PerplexityBot | Perplexity | Search / retrieval | PerplexityBot | You stop being surfaced and linked in Perplexity results. |
| Perplexity-User | Perplexity | User-triggered | Perplexity-User | Little. Perplexity says it "generally ignores robots.txt rules". |
| Google-Extended | Training control token (not a separate bot) | Google-Extended | No Gemini training, and no grounding in Gemini Apps or Vertex AI. No effect on Google Search. | |
| Applebot-Extended | Apple | Training control token (does not crawl) | Applebot-Extended | Out of Apple foundation-model training. Spotlight, Siri and Safari still show you. |
| CCBot | Common Crawl | Open web archive (we treat it as training) | CCBot | Future pages stay out of the public Common Crawl archive. |
Sources: OpenAI, Anthropic (updated 7 Apr 2026), Perplexity, Google (updated 14 Jul 2026), Apple, Common Crawl. All checked 22 September 2026.
Two rows deserve a second look.
Common Crawl describes its archive as open data "analyzable by anyone" and its FAQ doesn't list AI training among the uses. We still put CCBot with the training crawlers. That's our judgement: a bulk archive that anyone can download works like a training crawler for anyone who downloads it.
Google-Extended is the other one. It is the only "training" token here that also touches a live answer product, because Google's doc says it governs grounding in Gemini Apps as well as training. If Gemini app answers matter to you, blocking it has a cost.
Does blocking Google-Extended remove you from AI Overviews?
No. Google's crawler documentation says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search" (Google common crawlers).
AI Overviews and AI Mode are part of Search. Google's own page on AI features says "robots.txt directives for Googlebot is the control" for Search, and that you limit what shows with nosnippet, data-nosnippet, max-snippet or noindex (Google AI features doc, updated 10 Dec 2025).
So there is no robots.txt line that keeps you in Google's blue links but out of AI Overviews. People who add Google-Extended expecting that are blocking Gemini training and grounding, and nothing in Search changes.
Since 31 August 2026 there is a separate switch for that, and it lives in Search Console, not robots.txt. The Search generative AI control (Settings) lets a property opt out of AI Overviews, AI Mode and generative AI features in Discover. Google says it has rolled out to all websites worldwide, takes 1–2 days to apply, works per property rather than per page, and "isn't used as a ranking or inclusion signal" for the rest of Search. Our view: for most Indian businesses, opting out gives up visibility and gains nothing, so leave it on "include" unless you have a specific licensing reason. If AI Overviews are the goal, our post on ranking in AI Overviews is the next read.
One practical detail: Google says Google-Extended has no separate HTTP user-agent string. Crawling happens under the normal Google user agents. You will never see "Google-Extended" in your server logs, so don't go looking for it. The same goes for Applebot-Extended, which Apple says "does not crawl webpages".
Should you block AI crawlers? Our view
For most Indian businesses that sell something, our view is: block training, allow search. That's the middle row of the snippets below.
The reasoning is about what each crawler gives back. A training crawler takes your page and returns nothing you can measure. A search crawler takes your page and can return a citation with a link, which is a visit from someone already asking about your category. OpenAI's docs say the settings are independent: you can "allow OAI-SearchBot in order to appear in search results while disallowing GPTBot".
Kesar Lane fits this exactly. Its ingredient guides are the reason Perplexity might cite it for "kumkumadi oil for dry skin". Blocking PerplexityBot to protect the guides would remove the one thing they were written to earn.
Where we'd go further and block everything: a paywalled publisher whose article is the product, a research firm selling reports, or a site with licensing deals to protect. Where we'd allow everything, training included: a new brand with no audience that wants its name known to models as widely as possible. The mistake we most want to prevent is the cautious one: blocking a search crawler along with the training ones because both had "bot" in the name.
If being cited is the whole point, the robots.txt is only the permission slip. Getting chosen is a content job, which our guides on getting cited by ChatGPT and getting cited by Perplexity cover. When a brand wants that done end to end, it's what our AI SEO, AEO and GEO services team works on.
What robots.txt can't do
robots.txt is a request. RFC 9309, the standard that defines it, says plainly that it "is not a substitute for valid content security measures". A well-behaved crawler reads it and complies. Nothing forces one to.
That isn't hypothetical. On 4 August 2025 Cloudflare published tests saying Perplexity used "a generic browser intended to impersonate Google Chrome on macOS when their declared crawler was blocked" (Cloudflare). Perplexity disputed it, arguing the traffic was user-initiated fetching rather than crawling (Search Engine Land, 5 Aug 2025). Whichever side you believe, the lesson for a site owner is the same: if content must not be read, put it behind a login.
Three more limits people miss:
- Blocking is forward-looking. It doesn't remove pages a crawler already collected.
- Changes take time. OpenAI says it can take about 24 hours for its systems to adjust after you edit robots.txt.
- Blocking by IP can backfire. Anthropic warns that IP blocks "may not work correctly or persistently guarantee an opt-out" because they stop its bot reading your robots.txt at all.
How to set it up: decision flow, three snippets, log check
Work through these in order. The snippets use example.com and a typical D2C store's private paths; swap in your own.
- Decide per bucket, not per bot. Ask two questions. Do I want future models trained on my content? Do I want to be cited in AI answers? "No, yes" is Snippet B. "Yes, yes" is Snippet A. "No, no" is Snippet C.
- Snippet A: allow all crawlers, AI included. You don't need to name AI bots at all. Every crawler without its own group follows the
*group.# Snippet A: allow everything except private paths User-agent: * Disallow: /cart/ Disallow: /checkout/ Disallow: /my-account/ Sitemap: https://example.com/sitemap.xml - Snippet B: block training, allow search. The search bots and user fetchers have no group of their own here, so they fall back to
*and keep your normal rules. SeveralUser-agentlines can share one group.# Snippet B: no training, yes to AI search citations User-agent: * Disallow: /cart/ Disallow: /checkout/ Disallow: /my-account/ # Training crawlers and training control tokens User-agent: GPTBot User-agent: ClaudeBot User-agent: CCBot User-agent: Google-Extended User-agent: Applebot-Extended Disallow: / Sitemap: https://example.com/sitemap.xml - Snippet C: block all AI crawlers. This also removes you from ChatGPT search, Claude search and Perplexity citations. The two user fetchers are listed, but their vendors say they may not obey. OAI-AdsBot isn't listed because OpenAI says robots.txt doesn't control it.
# Snippet C: opt out of AI training and AI search User-agent: * Disallow: /cart/ Disallow: /checkout/ Disallow: /my-account/ User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: CCBot User-agent: Google-Extended User-agent: Applebot-Extended Disallow: / Sitemap: https://example.com/sitemap.xml - Don't give a search bot its own group unless you copy your private rules into it. Under RFC 9309 a crawler that finds a group with its own name ignores the
*group. AnOAI-SearchBotgroup with onlyAllow: /quietly opens your cart and account pages to it. - Check your server logs a few days later. Search the raw access log for the tokens that actually appear in traffic (not Google-Extended or Applebot-Extended, which never show up). On a Linux server with a combined-format log:
That prints IP, URL and status code. A blocked training bot should be requesting
BOTS="gptbot|oai-searchbot|chatgpt-user|claudebot" BOTS="$BOTS|claude-searchbot|claude-user" BOTS="$BOTS|perplexitybot|perplexity-user|ccbot" grep -Ei "$BOTS" access.log \ | awk '{print $1, $7, $9}' \ | sort | uniq -c | sort -rn | head -50/robots.txtand little else. - Confirm the IPs are real. Anyone can fake a user-agent string. Compare the IPs against each vendor's published list:
openai.com/gptbot.json,openai.com/searchbot.json,openai.com/chatgpt-user.json,claude.com/crawling/bots.json,perplexity.com/perplexitybot.jsonandperplexity.com/perplexity-user.json. A "GPTBot" hit from an IP outside OpenAI's list isn't OpenAI. - Generate and test the file. Build the final version in our free robots.txt generator, upload it to your domain root, and open
/robots.txtin a browser to confirm the live file is the one you wrote. On WordPress, check that an SEO plugin isn't serving a virtual robots.txt over yours.
For Kesar Lane, that's Snippet B, a log check on day three, and nothing else. Six lines of text decide whether its ingredient guides train a model or get cited by one.
Frequently asked questions
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects content that may be used to train OpenAI's foundation models. OAI-SearchBot fetches pages so they can be surfaced in ChatGPT's search features. OpenAI treats them as independent settings, so you can block GPTBot and allow OAI-SearchBot to stay out of training while still appearing in ChatGPT search answers.
Will blocking AI crawlers hurt my Google rankings?
Blocking GPTBot, ClaudeBot, PerplexityBot or CCBot has no connection to Google. Blocking Google-Extended doesn't affect Google Search either: Google says it is not used for Search inclusion or as a ranking signal. Only blocking Googlebot itself removes you from Google Search. To leave AI Overviews and AI Mode but stay in normal results, use the Search generative AI control in Search Console instead.
How do I block AI crawlers in robots.txt?
Add a group listing each bot's user-agent token followed by Disallow: /. For example, User-agent: GPTBot and User-agent: ClaudeBot on separate lines, then Disallow: / below them. Several User-agent lines can share one group. Use the exact tokens from each vendor's documentation, and remember the change applies to future crawling only.
Do AI crawlers respect robots.txt?
OpenAI, Anthropic, Perplexity, Google and Common Crawl all document robots.txt controls for their crawlers. User-triggered fetchers are the exception: OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it. robots.txt is a request, not enforcement, so check your server logs to see what actually happens.
Is llms.txt a replacement for blocking AI crawlers in robots.txt?
No. robots.txt tells crawlers what they may fetch. llms.txt is a separate, optional file that points AI tools to your most useful content, and it doesn't block anything. If you want a crawler kept out, robots.txt (or a login) is the control.
Get cited by the AI answers your customers read
We audit which AI crawlers can reach you, fix the rules, and build the content ChatGPT, Perplexity and AI Overviews choose to cite.
Make Digital Hangover a preferred source
One tap tells Google to show more of our SEO and marketing coverage in your Top Stories.
