All crawlers

Cohere · Training

cohere-training-data-crawler

The name announces a function that no Cohere document confirms. Its policy page disclaims training crawls "at this time" and leaves the bot table populated with `N/A`, so the token's stated purpose comes entirely from the string itself. Whether this and `cohere-ai` are two live agents or one that superseded the other is undeterminable from primary sources: there is no deprecation note, no changelog entry and no naming history in Cohere's docs index, and guessing between those readings would be inventing a fact. Cohere does publish a changelog for updates to that page, which is where a real name would first appear if one ever does.

Operated by
Cohere
Purpose
Training

Collects pages into a corpus used to train models.

robots.txt token
cohere-training-data-crawler
Verification
User agent only

How to verify cohere-training-data-crawler

Cohere publishes no list of IP addresses for cohere-training-data-crawler, so nobody (us included) can prove that a request carrying this user agent really came from Cohere. The name is trivial to copy. Treat it as a claim the visitor is making about itself, not as an identity anyone has checked.

There is no range list to confirm. This entry was last reviewed against Cohere's own documentation on .

User agent

Cohere has not published a full user-agent string for cohere-training-data-crawler. Requests are identified by the cohere-training-data-crawler product token appearing in the User-Agent header; we match that token rather than a whole string, because the rest of the header varies and matching it would miss real traffic.

robots.txt for cohere-training-data-crawler

Nothing token-specific exists. Cohere's general policy (that it "require[s] that crawlers be designed to respect `robots.txt`") is the whole of the written record, illustrated with the placeholder name `Coherebot` rather than this string. A `Disallow` written against `cohere-training-data-crawler` therefore rests on third-party reports of the token, not on any commitment Cohere has made about it.

Block

User-agent: cohere-training-data-crawler
Disallow: /

Allow

User-agent: cohere-training-data-crawler
Allow: /

robots.txt is a request, not an enforcement mechanism. It is honoured by convention, and a crawler that ignores it is stopped at your edge, not in a text file.

What blocking cohere-training-data-crawler costs you

There is no counterweight to place on the scale. A crawler that exists to assemble training data returns no visitors by construction (it reads once and sends nobody back), and Cohere operates no reader-facing surface where corpus membership could later put your name in front of someone. What allowing it risks is content entering a commercial training set with no attribution, no citation and no visibility into which models it reached. For an operator whose rule is to refuse training crawlers and permit retrieval ones, this token takes the least deliberation of anything in the directory.

Vendor documentation

Cohere does not publish documentation for cohere-training-data-crawler that we could find. Everything on this page comes from what they do publish elsewhere and from observed behaviour, so treat it accordingly.

Other Cohere tokens we track