The AI Crawlers Split Into Three. Most robots.txt Files Still Name One.
The file still returns 200. It stopped being true in 2024.
On 24 August 2026 I ran a full reconciliation of sagentix.ca/robots.txt against what the AI providers currently document — all twelve vendors the file names, page by page: OpenAI, Anthropic, Perplexity, Google, Bing, DuckDuckGo, Common Crawl, Apple, ByteDance, Meta, Amazon and Yandex. What came back was not a problem with the file. It was a structural change in how the providers work, and the file was the instrument that made it visible.
The file named 20 AI agents, every one Allow — 21 User-Agent blocks counting the wildcard. GPTBot, ClaudeBot, PerplexityBot, Googlebot: a deliberate, well-built file, and a better one than most sites carry. It carried names from an earlier generation of the vendor documentation, and it named anthropic-ai and Claude-Web — two strings that appear in none of Anthropic's current documentation — while naming neither Claude-User nor Claude-SearchBot, which did not exist when it was written.
Nothing was blocked. No monitor fired, the file returned 200, and it parsed cleanly. That is the whole finding: a correct file and a current file are different things, and no instrument on your side of the wire can tell them apart.
The spine: the providers split one crawler into three — one that trains, one that fetches when a user asks, one that indexes for search. A file written before that split names the training agent and misses the two that decide whether an assistant can reach and cite you. Nothing breaks when it happens, so nothing tells you.
What actually changed
The providers split one crawler into three, and they did it for a reason that makes sense once you see it: the three jobs have genuinely different consequences for a site owner, so they gave each one its own control.
Anthropic's documentation is the clearest statement of the pattern. It describes "the three robots that Anthropic uses" and separates them by function (Anthropic, 2026):
- ClaudeBot collects web content that may contribute to model training.
- Claude-User runs when a person asks Claude a question. Anthropic's own description of what disabling it does is the important sentence: it "prevents our system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search" (Anthropic, 2026).
- Claude-SearchBot indexes content for search-result quality.
Amazon documents the same three-way split — Amazonbot, Amzn-SearchBot, Amzn-User — and is explicit that they are separate controls: "Each user agent setting is independent of the others" (Amazon, 2026). Amzn-User "supports user actions, such as responding to Alexa queries that require up-to-date information" (Amazon, 2026).
Meta documents five crawlers rather than three, and the extra ones are the agentic and advertising cases. Meta-WebIndexer "navigates the web to improve Meta AI search result quality for users"; Meta-ExternalFetcher "fetches individual links at a user's request and supports product functions such as evaluating and improving agentic AI capabilities" (Meta, 2026).
OpenAI runs GPTBot for training, OAI-SearchBot for search and ChatGPT-User for user-initiated fetches, plus OAI-AdsBot for ad landing pages (OpenAI, 2026) — but its answer on the middle one is not Anthropic's, and the difference is the most useful thing on the page. OpenAI states that "ChatGPT-User is not used for crawling the web in an automatic fashion. Because these actions are initiated by a user, robots.txt rules may not apply" (OpenAI, 2026). For the search side it points owners elsewhere, recommending OAI-SearchBot as the agent to use in robots.txt for managing search opt-outs and automatic crawl (OpenAI, 2026). Anthropic presents Claude-User as an agent your file does govern, and warns that disabling it reduces your visibility. Same three jobs; two vendors, two different answers to whether your file governs the second one.
Perplexity runs PerplexityBot and Perplexity-User (Perplexity, 2026). DuckDuckGo documents DuckAssistBot beside DuckDuckBot (DuckDuckGo, 2026a, 2026b). Apple runs Applebot and Applebot-Extended (Apple, 2026).
Three jobs. Train, fetch-for-a-user, index-for-search. Once you know the shape, you can read any vendor's page in thirty seconds and say which of the three you have named.
Three weeks later, the three jobs became three switches
On 15 September 2026 Cloudflare shipped controls built on the same three-way split — and named the case this post's advice does not reach.
The taxonomy arrived independently. Cloudflare "classifies bots by behavior" and offers three of those behaviours as controls: Search, "crawling to build a search index"; Training, "crawling to train or fine-tune a model"; and Agent, "user-directed agents visiting a page on behalf of a human, such as chat fetch bots and browser-use agents" (Becker, 2026). Train, fetch-for-a-user, index-for-search — one layer down the stack, arrived at from the traffic rather than from the documentation. When the vendors and the network sitting in front of them sort the same requests into the same three boxes, the boxes are probably real.
The same post puts a number on how differently the three jobs are wanted. Fewer than 1% of Cloudflare sites block search crawlers, while 17% "choose to enable some mechanism to block training" (Becker, 2026). Almost nobody is trying to disappear. They are trying to separate two things that one file often cannot.
Which brings the complication, and it bounds what the method in this post can do for you. Some of the largest crawlers do two jobs under one name: "A mixed-use crawler is a single crawler doing both Search and Training" (Becker, 2026). Applebot, Bingbot and Googlebot are the named cases. For those three there is no pair of User-agent lines that separates training from search, because there is no second name to address — refuse one and you refuse the other. Reading your file against the three jobs reaches the vendors that split their crawlers (Amazon, Anthropic, Meta, Mistral, OpenAI) and runs out at the three that did not.
Cloudflare states the limit more bluntly than I did: "A robots.txt directive alone cannot solve this problem. Anyone can publish one, but it cannot identify who is crawling, determine why they are crawling, or stop a crawler that ignores it" (Becker, 2026). A file is a request addressed to a name. It cannot verify the name, and for a mixed-use crawler the name does not carry the distinction you want to draw.
The middle job stays the unsettled one, and four parties now give four answers. OpenAI says robots.txt rules may not apply to a user-initiated fetch (OpenAI, 2026). Anthropic presents Claude-User as an agent your file does govern, and warns that disabling it costs you visibility (Anthropic, 2026). Meta is the blunt one: because Meta-ExternalFetcher acts on a user's request, "this crawler may bypass robots.txt rules" (Meta, 2026). Cloudflare's answer is that the question has no instrument yet — "the Internet does not yet have a well-established directive for expressing Disallow preferences to agents" — so it ships no Disallow setting for Agents and points at the emerging ai-prefs standard (Becker, 2026). The job that runs when a buyer asks an assistant about you has the least settled machinery and the most commercial consequence.
The misread
Most people who notice this will reach the same conclusion I first did: add the missing names to the allow-list. That is the misread, and it matters because it points you at the wrong file.
If your robots.txt is a permissive file — User-agent: * with Allow: / and a handful of private paths disallowed — then the agents you failed to name were never blocked. The wildcard already granted them exactly what the named blocks granted. Mine was one of these. The practical damage was zero. What I had was a maintenance failure wearing the costume of a technical one.
The file that actually costs you something is the one with a Disallow in it.
Consider a site that decided, reasonably, to keep its content out of AI training while staying visible in AI search. In 2024 that was one line per vendor: disallow ClaudeBot, disallow GPTBot, leave everything else alone. Today that same file blocks training and permits both retrieval agents — which is what its author wanted, by luck.
Now reverse it. A site that wanted to stop all AI traffic wrote Disallow: / for ClaudeBot and GPTBot and considered the job done. That file now blocks training and permits Claude-User, Claude-SearchBot, ChatGPT-User and OAI-SearchBot — the three-way split moved two of the three jobs outside the rule its author wrote. The intent and the instruction have come apart, silently, and the file still looks deliberate.
Both sites have a file that reads as considered. Only one of them still says what its author meant. Neither owner has any way to tell from the outside, because a robots.txt never reports what it failed to match.
Google is the reason you cannot do this by log inspection
The obvious workaround is to skip the documentation and read your access logs: whatever user-agents show up, name those. That fails on Google, and Google is not a small exception.
Google's crawler documentation states plainly that "Google-Extended doesn't have a separate HTTP request user agent string" (Google, 2026a). It is a robots.txt control token and nothing else — crawling happens under existing Google user-agents, and the token governs whether the content is used for "training future generations of Gemini models" and for grounding in Gemini Apps and the Vertex AI API (Google, 2026a). It will never appear in a log file, because nothing ever sends it.
So the log-inspection method returns a list that is missing the single token controlling whether Google's AI products may use your content. A method that looks empirical, gives a confident answer, and cannot see the thing you were checking for is worse than no method — it retires the question.
The same trap sits one layer down, and it caught me. Google-Extended does not appear on Google's crawlers overview page at all — it is documented on the common crawlers page. The first version of this post said the special-case crawlers page, which I inferred and never opened; that page has zero mentions of it. Read only the overview, as any reasonable person would, and the evidence says the token has been retired: a fetch that succeeded, an answer that was confident, and nothing anywhere to signal the gap. I made exactly that error inside the paragraph warning about it, which is the most useful evidence I can offer that the check has to be a named set of pages rather than a search.
What I found when I checked all of them
The interesting question was whether Anthropic was an outlier. It is not: five vendors document a genuine three-way split — Anthropic, OpenAI, Amazon, Meta and Mistral. A separate and narrower question is where my own file had fallen behind, and that was four vendors. The two counts are not the same measurement, and the first version of this post presented them as one. Mistral is the cleanest of the five, because its page says what each token is not for: MistralAI-Training is "not used for search indexing or to answer live user queries", MistralAI-User is "not used for crawling the web in any automatic fashion, nor to crawl content for generative AI training", and MistralAI-Index is for "indexing purposes only" and "not used for generative AI training of any kind" (Mistral AI, 2026). Here is where my own file had fallen behind:
| Vendor | Documented as at 2026-08-24 | My file named |
|---|---|---|
| Amazon | Amazonbot, Amzn-SearchBot, Amzn-User |
Amazonbot |
| Meta | Meta-ExternalAgent, Meta-ExternalFetcher, Meta-WebIndexer, Meta-ExternalAds, facebookexternalhit |
Meta-ExternalAgent, plus FacebookBot |
| DuckDuckGo | DuckDuckBot, DuckAssistBot |
DuckDuckBot |
Googlebot, Google-Extended, GoogleOther, Google-CloudVertexBot (common crawlers page); Google-Agent, Google-GeminiNotebook (user-triggered fetchers page; Google, 2026b) — a selection, not Google's full set |
Googlebot, Google-Extended |
The pattern is not uniform, and saying so matters more than the tidy version. On Amazon the name carried forward was the training crawler and two retrieval agents were absent. Google is the reverse: Google-Extended, the training control, was named, and the user-triggered fetchers were not. DuckDuckGo documents no training crawler at all — DuckDuckBot is a search crawler, and of DuckAssistBot the company says "This data is not used in any way to train AI models" (DuckDuckGo, 2026b). What is common to all four is a gap between what the vendor documents and what the file named; what fills the gap differs by vendor.
FacebookBot is worth its own sentence. It appears nowhere in Meta's current crawler documentation, which lists five names and does not include it. It does exactly what anthropic-ai and Claude-Web do in a 2024 file: occupy a line, match nothing, and make the file look current.
One name on the list I got wrong, and the way I got it wrong is the point of this post. I wrote that ByteDance publishes no crawler page I could reach: developer.bytedance.com returns 404 and bytespider.bytedance.com does not resolve, both of which are true. The conclusion does not follow. ByteDance documents Bytespider on its Toutiao webmaster platform, at zhanzhang.toutiao.com/docs/intro/26899, which returns 200 from North America — and describes a search crawler: crawl the page, process it, serve retrieval. The word 训练, training, appears nowhere on it. The industry calls Bytespider an AI-training crawler; ByteDance's own and only documentation does not (ByteDance, 2026). I scoped the search to one hostname family and read absence as a finding, which is the exact failure the section above warns about.
Why this decays and nothing tells you
Every failure mode here shares one property: the artifact keeps working while it stops being true.
A blocked crawler produces symptoms. A stale user-agent string produces nothing at all. It matches no requests, appears in no log, triggers no alert, and leaves a file that reads as carefully maintained. The only instrument that detects it is re-reading the vendor's documentation and comparing — and that is a task nobody schedules, because there is no event that prompts it.
The vendor documentation moves too. Of those twelve pages, three had relocated: Google's crawler docs to a new path, Perplexity's to a different section, Anthropic's from support.anthropic.com to support.claude.com. All three still resolve through redirects, so a link check passes and a bookmark still works. The agent names had not changed in those three cases — but a process that treats "the link still works" as "the content is current" would not have known either way.
Two of the pages defeat a naive fetch, by different mechanisms. Bing's crawler page returns a JavaScript shell whose extractable text is one line — the page title — although the raw response is 125 KB and contains bingbot forty times inside a script payload (Microsoft, 2026). Meta's returns HTTP 400 to a browser user-agent, reproducibly, and the cause is not rendering: send a Googlebot user-agent and the same URL returns 324 KB of server-rendered HTML with all five crawler names in it. Grep the extracted text of either result for a user-agent string and you get zero matches, which reads exactly like this vendor documents no crawlers. It is not a small distinction. A failed fetch and an empty subject produce identical output, and only one of them is a finding.
The newer failure: the file changes while you do nothing
Everything above describes a file that stood still while the web moved. The same Cloudflare release introduces the opposite case, and it belongs in the check: a robots.txt that changes without its owner editing it.
Managed robots.txt does not advise. It writes. Where a site already has a file, "Cloudflare will prepend our managed robots.txt before your existing robots.txt, combining both into a single response"; where a site has none, "Cloudflare creates a new file with managed Disallow rules for known AI crawlers and serves it for you" (Cloudflare, 2026). The new Training control works the same way: Disallow AI Training "is named for the Disallow: directive it publishes in your robots.txt", and Bot Preference Sync "publishes the applicable no-training preference in robots.txt" (Becker, 2026).
So the file a crawler reads at your domain is your file plus whatever the network in front of it prepended. Reading the copy in your repository no longer tells you what you publish. That is a one-line change to the method — fetch the served file over HTTPS, the way a crawler does, and read that one.
The mechanism is mid-replacement too. "Managed Robots.txt will be deprecated in favor of Bot Preference Sync. Customers who enabled Managed Robots.txt will migrate to the new system" (Becker, 2026). That migration runs on Cloudflare's timetable, not the site owner's.
One setting changed meaning on the way through, which is this post's thesis wearing a dashboard instead of a text file. Block and "Block on pages with ads" "previously did not apply to mixed-use crawlers because blocking them could also affect search discoverability"; they now "apply to all training crawlers, including mixed-use crawlers" (Becker, 2026). No live behaviour flipped — a Training setting of Block migrates to Disallow AI Training, which preserves the practical effect (Becker, 2026). But the word is not the word it was. A runbook that says set Training to Block now points somewhere else than it did on 14 September, and Cloudflare is explicit about where: "It will stop Applebot, Bingbot, and Googlebot from reaching your site — search included" (Becker, 2026).
The sharpest case sits in the same announcement. Microsoft's robots.txt-level no-training support is still being built, targeted for early 2027, so "until that support launches, selecting Disallow AI Training will not automatically convey a no-training preference to Bing through robots.txt" (Becker, 2026). A control named for an outcome, correctly configured, reporting success — and not yet producing that outcome at one of the three crawlers it exists to address. Cloudflare publishes the caveat in the launch post, which is the right way to ship it. It is still a name describing a destination rather than a state.
How Sagentix handles this
The check runs as a registry and a schedule rather than an intention. Every agent string, its role, the vendor page that documents it, and a canary token that proves a fetch actually worked — because a fetch that fails must never read as an absence. A weekly job re-reads all thirteen vendor pages and reports drift: agents no longer documented, new agent-shaped names on the page, documentation that moved, and any gap between what the vendors document and what the site's own file says. It is silent on a clean week and notifies on a change.
It earns its keep on weeks like this one. Re-run on 16 September 2026 against the very file this post corrected, it returned drift. Google's user-triggered fetchers page now lists Google-NotebookLM as a "Former agent (supported until August 2026)" and documents Google-GeminiNotebook in its place (Google, 2026b). The registry still held the retired name, so it read the current agent as unrecognised and asked for the dead one — the instrument for catching stale agent strings had gone stale itself, twenty-three days after this post went up. The same run found Mistral named in the site's file and in this article but absent from the registry, so nothing was watching the vendor described here as the cleanest of the five. The same run found Meta's crawler page had moved from /docs/ to /documentation/, and the old path still redirects — so no link check would ever have reported it. All three are corrected as of 16 September 2026. That is the argument for the schedule, and the schedule makes it better than any assertion of mine could.
This is the same evidence discipline that runs on every client deliverable — 1,425 curated IP artifacts behind an 18-check quality gate, a 6–8 week engagement, CA$4.5K–$45K, with a Phase 1 money-back guarantee (subject to terms). The Phase 09 Digital Audit treats AI-retrieval configuration as a dated observation with a re-check interval, not a one-time fix, for exactly the reason this post exists Sagentix Phase 09 Digital Audit, 2026.
The general lesson is the one worth taking, and it is not about robots.txt. Expertise produces a correct artifact on the day it is written. Nothing about expertise keeps that artifact correct while the thing it describes moves underneath it. Only a re-read on a schedule does that — which is why the interesting deliverable here was never the corrected file. It was the weekly job that will catch the next change, and the one after.
Three things you can do with this
Read your own file against the three jobs, not the vendor names. Open yourdomain.com/robots.txt on your phone — the served file, not the copy in your repository, because the network in front of you may be prepending lines to it. For each vendor you have named, ask which of train / fetch-for-a-user / index-for-search you have addressed. If you named one string, you probably addressed one job. For Applebot, Bingbot and Googlebot, ask a different question: those are mixed-use, so the separation you want is a setting at the edge rather than a line in the file (Becker, 2026). Twenty minutes, no tools, no subscription. If you have any Disallow line at all, do this today rather than this quarter — that is the file where intent and instruction come apart.
Put a re-read on the calendar, quarterly. The names change on the vendors' schedule, not yours, and nothing about the decay is observable from your side. A recurring reminder to re-read the vendor pages, named one by one rather than counted, beats any amount of care taken once. Fetch each vendor's page with a real browser user-agent, and confirm each fetch returned actual content before concluding anything is missing.
Or have it audited as part of a broader diagnostic, which is what the Phase 09 Digital Audit does — with the standing caveat that the audit is only worth what its re-check interval is worth. If you take one idea from this piece and never speak to me, take the first one; it is where nearly all the value sits.
A last note on scope: Sagentix advises on go-to-market strategy. I do not deliver certification, audit or legal services, and nothing here is a route into selling any.
References
- Amazon. (2026). Amazonbot. Amazon Developer. https://developer.amazon.com/amazonbot
- Anthropic. (2026, April 7). Does Anthropic crawl data from the web, and how can site owners block the crawler? Anthropic Help Center. https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Apple. (2026). About Applebot. Apple Support. https://support.apple.com/en-us/119829
- Becker, B. (2026, September 15). Have it both ways: stay discoverable in search while disallowing AI training. Cloudflare Blog. https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/
- ByteDance. (2026). 关于Bytespider [About Bytespider]. Toutiao Webmaster Platform. https://zhanzhang.toutiao.com/docs/intro/26899
- Cloudflare. (2026). Managed robots.txt. Cloudflare Docs. https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/
- DuckDuckGo. (2026a). DuckDuckBot. DuckDuckGo Help Pages. https://duckduckgo.com/duckduckgo-help-pages/results/duckduckbot/
- DuckDuckGo. (2026b). DuckAssistBot. DuckDuckGo Help Pages. https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot/
- Google. (2026a). Google crawlers and fetchers: Google's common crawlers. Google Search Central. https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
- Google. (2026b). Google crawlers and fetchers: User-triggered fetchers. Google Search Central. https://developers.google.com/crawling/docs/crawlers-fetchers/google-user-triggered-fetchers
- Meta. (2026). Web crawlers. Meta for Developers. https://developers.facebook.com/docs/sharing/webmasters/web-crawlers (returns HTTP 400 to a browser user-agent; serves the full page to a Googlebot user-agent)
- Microsoft. (2026). Which crawlers does Bing use? Bing Webmaster Tools Help. https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0
- Mistral AI. (2026). Robots. Mistral AI Documentation. https://docs.mistral.ai/robots
- OpenAI. (2026). OpenAI bots. OpenAI Developers. https://developers.openai.com/api/docs/bots
- Perplexity. (2026). Perplexity crawlers. Perplexity Docs. https://docs.perplexity.ai/docs/resources/perplexity-crawlers
Subscribe + get the workbook
The Bottom-Up TAM / SAM / SOM Workbook — free with your subscription
An 11-page tactical workbook with fillable worksheets — NAICS lookup, three-filter SAM test, Bull/Base/Bear SOM, and the diligence cross-checks. Not published anywhere else. Then get evidence-backed analysis every other Tuesday. No spam. Unsubscribe anytime. See past issues.

Stéphane Raby
Founder & Principal — Sagentix Advisors
CMC | CISSP | P.Eng. | uOttawa Telfer Executive MBA — ranked #1 globally by CEO Magazine, 2023. 25+ years in technology strategy, cybersecurity, and management consulting.
Want This Evidence Applied to Your Market?
Phase 1 Market Intelligence starts at CA$4,500 with a money-back guarantee.