AI crawler
An AI crawler is an automated client that an AI company runs to fetch web pages, whether to train a model, build a search index or answer a question in real time. It names itself in its user agent and runs no JavaScript, so a browser analytics tag never sees it.
In Traceten: Traceten's directory lists 60 crawlers from 21 providers. Their hits reach Traceten from your server or Cloudflare, not from the page.
Answer fetch
An answer fetch is a request an AI assistant makes in real time to read a page while it answers a question someone just asked. ChatGPT-User, Claude-User and Perplexity-User are answer-fetching crawlers.
In Traceten: Traceten files these under Answers, one of four crawler categories. It is the closest observable sign that your page was read for a live answer.
Search index crawler
A search index crawler fetches pages continuously to build and refresh the index an AI search product answers from. OAI-SearchBot, Claude-SearchBot and PerplexityBot are search index crawlers.
In Traceten: Traceten files these under Search index, and its robots.txt checker reports each one apart from the same vendor's training crawler, such as OAI-SearchBot beside GPTBot.
Training crawler
A training crawler collects pages into a corpus used to train AI models. GPTBot, ClaudeBot and CCBot are training crawlers.
In Traceten: Traceten counts these under Training, so they never mix with the answer fetches that read a page for a live question.
robots.txt token
A robots.txt token is the name a crawler answers to on a robots.txt User-agent line. Most tokens match a crawler's user agent, but a few, such as Google-Extended and Applebot-Extended, are permission tokens that control how content is used and never send a request.
In Traceten: Because a permission token never fetches anything, Traceten marks a live request that carries one as spoofed.
Spoofed crawler
A spoofed crawler is a request that claims a crawler's user agent but comes from outside the network its vendor publishes. Anyone can type GPTBot into a user agent, so the claim alone proves nothing.
In Traceten: Traceten marks a hit spoofed when the claimed vendor publishes its IP ranges and the address falls outside them, or when it carries a permission token that never fetches, and keeps those hits apart from genuine ones.
Crawler verification
Crawler verification is checking that a request claiming to be a crawler really comes from its vendor. The strongest checks match the address against the vendor's published IP ranges or confirm it with reverse DNS.
In Traceten: Every crawl in Traceten records its tier and a confidence: IP range 0.99, reverse DNS 0.95, ASN match 0.85, user agent only 0.60 and spoofed 0. User agent only is normal for a vendor that publishes no ranges.