xAI · Other AI
GrokBot
Bulk web acquisition demonstrably happens somewhere upstream of the model: xAI states that Grok was pre-trained on a large corpus of publicly available information, including raw web page data, metadata extracts and text extracts from the Internet. Whether xAI collects that itself with a first-party crawler, buys it, or takes it from a third-party corpus is never said, and no agent name is ever attached to the step. Read this token as an unattributed candidate for that role rather than a documented one. The user-agent string widely copied for it traces to secondary directories only, and the site its trailing URL points at does not mention the bot, so it fails the usual convention that the URL in a user agent explains the agent.
- Operated by
- xAI
- Purpose
- Other AI
- robots.txt token
- GrokBot
- Verification
- User agent only
AI-adjacent traffic that does not fit the three purposes above.
How to verify GrokBot
xAI publishes no list of IP addresses for GrokBot, so nobody (us included) can prove that a request carrying this user agent really came from xAI. The name is trivial to copy. Treat it as a claim the visitor is making about itself, not as an identity anyone has checked.
There is no range list to confirm. This entry was last reviewed against xAI's own documentation on .
User agent
xAI has not published a full user-agent string for GrokBot. Requests are identified by the GrokBot product token appearing in the User-Agent header; we match that token rather than a whole string, because the rest of the header varies and matching it would miss real traffic.
robots.txt for GrokBot
Block
User-agent: GrokBot
Disallow: /Allow
User-agent: GrokBot
Allow: /robots.txt is a request, not an enforcement mechanism. It is honoured by convention, and a crawler that ignores it is stopped at your edge, not in a text file.
What blocking GrokBot costs you
Corpus inclusion sends no visitors and produces no citation, so measured on referral metrics a disallow costs nothing you can observe. What it costs is verification in the other direction: xAI publishes no dataset and no opt-out registry, so unlike an openly released corpus there is no way to check afterwards whether the rule took effect, or what had already been taken before you wrote it. One asymmetry is worth printing plainly. x.ai's own robots.txt carries a draft Content-Signal line asserting ai-train=no over its content, alongside named disallows for other vendors' training crawlers, a preference xAI declares for itself while documenting no commitment to honour the equivalent signal on your site.
Vendor documentation
xAI does not publish documentation for GrokBot that we could find. Everything on this page comes from what they do publish elsewhere and from observed behaviour, so treat it accordingly.