Meta · Training
meta-externalagent
Model pretraining and product indexing are fed by the same sweep here, a wider mandate than any sibling carries: Meta's stated use cases are “training foundation AI models or improving products by indexing content directly.” No per-request trigger exists: it arrives on Meta's schedule rather than a visitor's, which is exactly what separates it from meta-externalfetcher's single on-demand pull and meta-webindexer's freshness crawl. Meta's wording has drifted between revisions: the January 2026 page said “training AI models”; the current one says “training foundation AI models.” Verification is a dead end and recently became a worse one: through the 7 January 2026 capture the page carried a Crawler IPs section pointing at an RADB route dump for AS32934, and the 21 May 2026 revision removed it, leaving the user-agent string as the only thing to match on.
- Operated by
- Meta
- Purpose
- Training
- robots.txt token
- meta-externalagent
- Verification
- Network origin only
Collects pages into a corpus used to train models.
How to verify meta-externalagent
We hold no list of IP addresses that Meta publishes itself for meta-externalagent. What we hold instead is 574 prefixes that a third party (a public routing record, not Meta) reports Meta's network announcing. A request from inside one of those tells you where it came from, not who sent it: anything else on that network can send the same packet, and nothing we hold ties meta-externalagent to those addresses beyond the network they sit on. A request from OUTSIDE them is not thereby a fake either: this describes one network's announcements, not every address Meta can crawl from. Corroboration, not proof.
We hold no range list Meta publishes itself for meta-externalagent, so there is no such fetch to confirm. This entry was last reviewed against Meta's own documentation on .
User agent
Meta publishes this user agent for meta-externalagent. Match on the meta-externalagent product token rather than the whole string: vendors revise the surrounding version and URL fragments without notice.
meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)robots.txt for meta-externalagent
Yes, with no carve-out: Meta's robots.txt section uses this exact token as its worked example, and the two bypass exceptions it names apply to meta-externalfetcher and facebookexternalhit, not to this one. Meta warns that robots.txt is cached for up to 24 hours, so a new rule takes up to a day to bite.
Block
User-agent: meta-externalagent
Disallow: /Allow
User-agent: meta-externalagent
Allow: /robots.txt is a request, not an enforcement mechanism. It is honoured by convention, and a crawler that ignores it is stopped at your edge, not in a text file.
What blocking meta-externalagent costs you
Nothing a visitor can see goes dark. No preview breaks, no citation vanishes, no ranking moves; the entire cost sits in whether your text is available to Meta's model training and to the direct-indexing pipeline bundled under the same token. Because Meta never enumerates which product indexes are built from that crawl, a blanket disallow forfeits an unnamed set rather than a listed one, and that ambiguity, not the training question, is the only real reason to hesitate.