AI Crawler Directory
Every AI crawler token we track, with what it actually does, the robots.txt directive to allow or block it, and an honest answer to the question most directories skip: can anyone prove this crawler is really who it says it is?
- User agent only
Alibaba
Qwen-User
The `-User` suffix is the only thing anyone has to go on here, and a suffix is not documentation.
Answers
- User agent only
Alibaba
QwenBot
Alibaba confirms the activity and never names the actor.
Training
- User agent only
Alibaba
TongyiBot
Two questions hide inside this token and only one of them can be settled from primary sources.
Other AI
- User agent only
Allen AI
AI2Bot
What this collects becomes a published artefact rather than a private asset.
Training
- IP verified
Amazon
Amazonbot
Of Amazon's three tokens this is the only one carrying no training carve-out, and the omission is the point: Amzn-SearchBot and Amzn-User both state they do not crawl content for generative AI model training, while this one is described as improving Amazon's products and services and "may be used to train Amazon AI models" (note the hedge).
Training
- IP verified
Amazon
Amzn-SearchBot
Amazon names a concrete retrieval surface for this one: permitting it makes your content "eligible to appear in search experiences such as Alexa," and the load-bearing sentence is the disclaimer that it "does not crawl content for generative AI model training." Where Amazonbot's crawl may reach a model's weights, this crawl builds an index that gets queried, and it builds it ahead of time, which is what separates it from Amzn-User's fetch at the moment a question is asked.
Search index
- IP verified
Amazon
Amzn-User
A customer's live question is the trigger: Amazon describes fetching "live information from the web to provide accurate answers on the user's behalf," typically an Alexa query that requires up-to-date information.
Answers
- IP verified
Anthropic
Claude-SearchBot
Index quality is the whole remit: Anthropic describes this agent as navigating the web to improve search result quality, analysing content specifically to enhance the relevance and accuracy of search responses.
Search index
- IP verified
Anthropic
Claude-User
A person asked Claude something, and Claude went to fetch your page on their behalf.
Answers
- IP verified
Anthropic
ClaudeBot
Bulk collection of public web content that could potentially contribute to Anthropic's model training.
Training
- IP verified
Apple
Applebot
One crawler, three downstream uses, each with its own separate off switch, which is why a blanket disallow here is almost always the wrong instrument.
Training
- Not a crawler
Apple
Applebot-Extended
There is nothing here to catch in a log file, and that fact is worth more than a verification badge would be: Apple states plainly that "Applebot-Extended does not crawl webpages" and that it "is only used to determine how to use the data crawled by the Applebot user agent." Because no request is ever issued under this name, it sits at no rung of the verification ladder at all (not IP range, not reverse DNS, not even the user-agent-only floor), so any inbound hit carrying this string is spoofed by definition, with no benign explanation available.
Training
- Vendor-documented
Baidu
Baiduspider
Reverse DNS is not a waypoint on the road to an IP list here: it is where verification stops, permanently, because Baidu has twice refused in writing to publish one:「Baiduspider的IP池是不断变动的,我们无法提供IP全集」("Baiduspider's IP pool changes constantly; we cannot provide the complete set of IPs"), and three years later「IP地址范围动态变化不固定,我们无法对外公布」("the IP address ranges change dynamically and are not fixed; we cannot publish them externally").
Search index
- User agent only
Baidu
ERNIEBot
Nine rows is how many crawler user agents Baidu enumerates in its webmaster FAQ (`Baiduspider` for PC, mobile and other search, then `-image`, `-video`, `-news`, `-favo`, `-cpro` and `-ads`), and this token appears in none of them.
Training
- User agent only
Baidu
YiyanBot
The product this token is named after no longer carries the name.
Other AI
- Vendor-documented
ByteDance
Bytespider
Widely filed as a training crawler, this is the one token in this group with real first-party documentation, and that documentation describes a search engine.
Search index
- User agent only
ByteDance
Doubaobot
Doubao (豆包) is by user count among the largest consumer AI assistants in China, which makes the documentation vacuum around its crawler far more conspicuous than the same silence around the smaller names.
Other AI
- User agent only
ByteDance
TikTokSpider
The most informative fact about this token belongs to its sibling.
Search index
- User agent only
Cohere
cohere-ai
No primary source establishes what this does, and the vendor page that ought to comes closer to denying it exists: "We do not use Cohere bots or user agents for the purpose of crawling or scraping web content to train generative AI foundation models at this time." The accompanying bot table has one row, reading `N/A | N/A`, and Cohere undertakes to "identify those bots and/or user agents in the table below" should that change.
Training
- User agent only
Cohere
cohere-training-data-crawler
The name announces a function that no Cohere document confirms.
Training
- IP verified
Common Crawl
CCBot
No product sits at the other end of this one: no search box, no assistant, no answer card.
Training
- User agent only
DeepSeek
DeepSeekBot
The collection is admitted; the collector is not.
Training
- IP verified
DuckDuckGo
DuckAssistBot
Two claims on DuckDuckGo's page carry this entry, one negative and one positive.
Answers
- IP verified
Google
Google-Agent
Something a person asked an AI agent to do (book the thing, compare the prices, fill in the form) is what puts this on your server.
Answers
- IP verified
Google
Google-CloudVertexBot
The only token here that a site owner asks for.
Training
- Not a crawler
Google
Google-Extended
Every other entry in this directory describes a program that makes requests.
Training
- IP verified
Google
Google-GeminiNotebook
The request is downstream of a click.
Answers
- IP verified
Google
Google-InspectionTool
You cause this one yourself.
Search index
- IP verified
Google
Google-NotebookLM
Read this entry as a migration notice rather than a crawler description.
Answers
- IP verified
Google
Google-Read-Aloud
The trigger is somebody pressing play.
Answers
- IP verified
Google
Googlebot
The oldest and broadest thing in this directory, and the reason the robots.txt convention exists at all.
Search index
- IP verified
Google
GoogleOther
Deliberately unattached to any product, and that is the entire design: Google states that preferences addressed to this token do not affect any specific product, and calls it the generic crawler various product teams may use for fetching publicly accessible content.
Other AI
- Network origin only
Meta
facebookexternalhit
A URL shared into a Facebook post, a Messenger thread or a Facebook social plugin brings this within seconds to build the card the recipient will see; Meta says it “gathers, caches, and displays” the title, description and thumbnail.
Other AI
- Network origin only
Meta
meta-externalads
One sentence is the entire published description: crawling “for use cases such as improving advertising and other business-related products and services.” There is no trigger, no cadence, no named product and no stated relationship to anything a user does, so what can be said with confidence is negative: it is not the link-preview path, not the Meta AI citation index, and not the model-training sweep.
Other AI
- Network origin only
Meta
meta-externalagent
Model pretraining and product indexing are fed by the same sweep here, a wider mandate than any sibling carries: Meta's stated use cases are “training foundation AI models or improving products by indexing content directly.” No per-request trigger exists: it arrives on Meta's schedule rather than a visitor's, which is exactly what separates it from meta-externalfetcher's single on-demand pull and meta-webindexer's freshness crawl.
Training
- Network origin only
Meta
meta-externalfetcher
Ask Meta AI a question and this is the request that lands on the one page needed to answer it.
Answers
- Network origin only
Meta
meta-webindexer
Citations with links are the stated payback, and this is the only Meta token whose documentation promises anything back at all: Meta says it “navigates the web to improve Meta AI search result quality,” analysing content “to enhance the relevance and accuracy of Meta AI.” That attribution surface is the concrete difference from meta-externalagent, which absorbs the same text into training with no link attached to it.
Search index
- IP verified
Microsoft
Bingbot
“Our standard crawler,” in Microsoft's words, handling “most of our crawling needs each day”: the general-purpose discovery and refresh crawl behind the Bing index, in desktop and mobile variants.
Search index
- Not sent as a user agent
Microsoft
msnbot
Dead as a crawler, alive as a verification artifact: that is the whole of it.
Search index
- IP verified
Mistral
MistralAI-Index
Scheduled sweeping is the job: Mistral calls it "automated crawling of the web for indexing purposes only" and says it "indexes content for Mistral search, which helps answer user questions in Vibe." Same destination as `MistralAI-User`, opposite timing: this one builds the corpus before anybody asks, where the user token fetches the single page a live question needs.
Search index
- User agent only
Mistral
MistralAI-Training
Dataset construction is the declared function: Mistral says it "crawls web content to help build datasets for training Mistral generative AI models," and rules out the other two roles its siblings hold: "This crawler is not used for search indexing or to answer live user queries in Vibe." What makes it worth a page of its own is the verification asymmetry inside one vendor's documentation.
Training
- IP verified
Mistral
MistralAI-User
Three tokens, three jobs, three directives: this is the one that fires while somebody waits.
Answers
- IP verified
Moonshot AI
Kimi-SearchBot
Relevance scoring at fetch time is what separates this agent from an archiver: Moonshot says it "analyzes pages for relevance and builds the search index" that Kimi's search features query, so its output is a ranked corpus consulted when someone asks a question, not a static pile of text.
Search index
- IP verified
Moonshot AI
Kimi-User
Someone in a Kimi session pastes a link and asks what it says; that is the whole trigger, and Moonshot's own examples are summarising a specific article or answering a question that needs live web retrieval.
Answers
- IP verified
Moonshot AI
KimiBot
Nothing this token fetches reaches a surface anyone will see.
Training
- IP verified
OpenAI
ChatGPT-User
Someone typed a question, or pasted your URL, and ChatGPT went and got the page while they waited: one request, one human, one moment.
Answers
- IP verified
OpenAI
GPTBot
Corpus collection, and nothing else: this is the token that decides whether your pages can end up inside a future GPT model's weights.
Training
- IP verified
OpenAI
OAI-AdsBot
Fires only against URLs somebody has already submitted as a ChatGPT ad landing page: OpenAI states it visits only pages submitted as ads.
Other AI
- IP verified
OpenAI
OAI-SearchBot
Everything ChatGPT quotes, summarises and links in its search answers comes out of the index this token builds.
Search index
- IP verified
Perplexity
Perplexity-User
Perplexity's public argument is that this is not a crawler at all: "User-driven agents only act when users make specific requests, and they only fetch the content needed to fulfill those requests.
Answers
- IP verified
Perplexity
PerplexityBot
Two things separate this token from its sibling: it chooses when to visit, and it honours robots.txt when it does.
Search index
- User agent only
xAI
Grok
The product name, not an agent token: every xAI page fetched uses this word to mean the assistant itself, and none binds it to a fetching program.
Other AI
- User agent only
xAI
Grok-DeepSearch
One prompt, tens of fetches: that is the shape a deep-research run leaves in a log, and the mode itself is documented rather than invented.
Answers
- User agent only
xAI
GrokBot
Bulk web acquisition demonstrably happens somewhere upstream of the model: xAI states that Grok was pre-trained on a large corpus of publicly available information, including raw web page data, metadata extracts and text extracts from the Internet.
Other AI
- User agent only
xAI
xAI-Bot
Nobody has established what this string does, and the useful version of this page says so rather than picking a side.
Other AI
- User agent only
xAI
xAI-Grok
Set this token beside GrokBot and almost nothing separates the two.
Other AI
- User agent only
xAI
xAI-SearchBot
Grok can ground an answer in a live lookup rather than in model memory: xAI's consumer FAQ confirms real-time web search as a product feature, and its developer docs confirm that fetched sources come back to the user as inline citations with full traceability.
Answers
- User agent only
xAI
xAI-Web-Crawler
Rules against this string are already deployed in production robots.txt files on the open web, which is the interesting fact about it, because its vendor has never acknowledged that it exists.
Other AI
- Vendor-documented
You.com
YouBot
Verification is where this one is genuinely unusual, and it is three methods deep: requests are signed under the Web Bot Auth standard with Ed25519 keys served from a live well-known signature directory, hostnames follow a documented reverse-DNS pattern with a worked forward-confirmation example, and legitimate traffic is stated to originate from a single published /24.
Search index
- User agent only
Zhipu AI
ChatGLM-Spider
A `+` URL inside a user agent is supposed to lead to the crawler's own account of itself, which makes the one in circulation for this token worth actually opening: `https://zhipu.ai/spider` serves no crawler policy.
Training