All crawlers

Cohere · Training

cohere-ai

No primary source establishes what this does, and the vendor page that ought to comes closer to denying it exists: "We do not use Cohere bots or user agents for the purpose of crawling or scraping web content to train generative AI foundation models at this time." The accompanying bot table has one row, reading `N/A | N/A`, and Cohere undertakes to "identify those bots and/or user agents in the table below" should that change. Verification is `user_agent_only` with no route upward, and Cohere effectively confirms the ceiling, warning that "other methods, such as blocking IP addresses, may not work reliably or ensure persistent opt-out, as they prevent the reading of a site's `robots.txt` file." No ranges, no signatures, no reverse DNS: a name in a header and nothing behind it.

Operated by
Cohere
Purpose
Training

Collects pages into a corpus used to train models.

robots.txt token
cohere-ai
Verification
User agent only

How to verify cohere-ai

Cohere publishes no list of IP addresses for cohere-ai, so nobody (us included) can prove that a request carrying this user agent really came from Cohere. The name is trivial to copy. Treat it as a claim the visitor is making about itself, not as an identity anyone has checked.

There is no range list to confirm. This entry was last reviewed against Cohere's own documentation on .

User agent

Cohere has not published a full user-agent string for cohere-ai. Requests are identified by the cohere-ai product token appearing in the User-Agent header; we match that token rather than a whole string, because the rest of the header varies and matching it would miss real traffic.

robots.txt for cohere-ai

Undocumented for this token, because Cohere names no tokens. The only applicable statement is a company-level one ("It is Cohere's policy to require that crawlers be designed to respect `robots.txt`"), which is a forward-looking commitment about crawlers Cohere might run, not a compliance claim about a deployed agent. The single user-agent string that appears anywhere on Cohere's policy page is `Coherebot`, used as a placeholder inside an example robots.txt block, and it matches neither registry token.

Block

User-agent: cohere-ai
Disallow: /

Allow

User-agent: cohere-ai
Allow: /

robots.txt is a request, not an enforcement mechanism. It is honoured by convention, and a crawler that ignores it is stopped at your edge, not in a text file.

What blocking cohere-ai costs you

Weigh the two sides and one of them is empty. Cohere sells models and APIs to enterprises and runs no consumer assistant or answer engine that cites sources back to publishers, so allowing this agent cannot return a visitor to you: there is no product on the far end with a citation slot to be listed in. That makes the trade-off one-sided in a way it is not for a vendor with an answer surface, where a disallow costs a real referral path. Treat that as our assessment of what Cohere's own materials describe and do not describe, rather than as something Cohere has stated.

Vendor documentation

Cohere does not publish documentation for cohere-ai that we could find. Everything on this page comes from what they do publish elsewhere and from observed behaviour, so treat it accordingly.

Other Cohere tokens we track