Cohere · Training
cohere-training-data-crawler
The name announces a function that no Cohere document confirms. Its policy page disclaims training crawls "at this time" and leaves the bot table populated with `N/A`, so the token's stated purpose comes entirely from the string itself. Whether this and `cohere-ai` are two live agents or one that superseded the other is undeterminable from primary sources: there is no deprecation note, no changelog entry and no naming history in Cohere's docs index, and guessing between those readings would be inventing a fact. Cohere does publish a changelog for updates to that page, which is where a real name would first appear if one ever does.
- Operated by
- Cohere
- Purpose
- Training
- robots.txt token
- cohere-training-data-crawler
- Verification
- User agent only
Collects pages into a corpus used to train models.
How to verify cohere-training-data-crawler
Cohere publishes no list of IP addresses for cohere-training-data-crawler, so nobody (us included) can prove that a request carrying this user agent really came from Cohere. The name is trivial to copy. Treat it as a claim the visitor is making about itself, not as an identity anyone has checked.
There is no range list to confirm. This entry was last reviewed against Cohere's own documentation on .
User agent
Cohere has not published a full user-agent string for cohere-training-data-crawler. Requests are identified by the cohere-training-data-crawler product token appearing in the User-Agent header; we match that token rather than a whole string, because the rest of the header varies and matching it would miss real traffic.
robots.txt for cohere-training-data-crawler
Nothing token-specific exists. Cohere's general policy (that it "require[s] that crawlers be designed to respect `robots.txt`") is the whole of the written record, illustrated with the placeholder name `Coherebot` rather than this string. A `Disallow` written against `cohere-training-data-crawler` therefore rests on third-party reports of the token, not on any commitment Cohere has made about it.
Block
User-agent: cohere-training-data-crawler
Disallow: /Allow
User-agent: cohere-training-data-crawler
Allow: /robots.txt is a request, not an enforcement mechanism. It is honoured by convention, and a crawler that ignores it is stopped at your edge, not in a text file.
What blocking cohere-training-data-crawler costs you
There is no counterweight to place on the scale. A crawler that exists to assemble training data returns no visitors by construction (it reads once and sends nobody back), and Cohere operates no reader-facing surface where corpus membership could later put your name in front of someone. What allowing it risks is content entering a commercial training set with no attribution, no citation and no visibility into which models it reached. For an operator whose rule is to refuse training crawlers and permit retrieval ones, this token takes the least deliberation of anything in the directory.
Vendor documentation
Cohere does not publish documentation for cohere-training-data-crawler that we could find. Everything on this page comes from what they do publish elsewhere and from observed behaviour, so treat it accordingly.