Skip to content
Blog

AI search

AI Crawlers on a Brand-New Site: 4,805 Hits and 14 Visits in 25 Days

We counted the AI crawlers on traceten.com for 25 days: 4,805 hits from 11 companies, 17 impostors, 14 AI-referred visits. What it shows, and what it does not.

Jay Patel10 min read

Between 5 and 29 September 2026, AI crawlers requested traceten.com 4,805 times. The hits came from 18 crawlers run by 11 companies, and OpenAI alone sent 2,186 of them. Another 17 requests used OpenAI's or Anthropic's name from addresses those companies do not publish. Over the same 25 days, AI assistants sent 14 visits.

This is our own marketing site in its first month of crawler tracking, measured with the product it describes. It is one young site with little history, so treat it as a case study and not an industry benchmark. What it can show is the shape of the thing: which crawlers turn up first, what they read, how many of them are who they say they are, and how far crawling runs ahead of the people it eventually sends.

How did we measure it?

traceten.com runs Traceten's own crawler middleware, the same @traceten/ai-crawl package a customer installs. Because a crawler never runs your tracking script, the count is taken on the server. Every request whose user agent matches a crawler in the AI crawler directory is reported, except requests for scripts, stylesheets, images and fonts. robots.txt, llms.txt and the sitemap always count. What one crawl hit is covers those rules in more detail.

A few things to know before reading the numbers:

  • The window. We queried 30 August to 29 September 2026, but crawl reporting on the site began on 5 September, so every crawl here falls in the 25 days from 5 to 29 September. All dates are UTC. We stopped at the last complete day and read the figures through Traceten's own API on 30 September.
  • Who counts as an AI crawler. The directory lists 60 crawlers from 21 companies. It includes Googlebot, because Google names robots.txt rules for Googlebot as the control for its AI features in Search, so Google's total below includes ordinary search crawling.
  • Verification. Every hit is checked against what its vendor publishes, as described in the verification runbook. A request that uses a vendor's name from an address outside that vendor's published ranges is marked spoofed and counted separately. The 17 spoofed requests are never part of the 4,805.
  • A floor, not a ceiling. The middleware reports in the background and never delays a page. A report that fails to arrive is simply not counted, so every figure here is a minimum.
  • No customer data. This is our marketing site. Its robots.txt allows every crawler, so nothing below was shaped by a block.

Who crawled the site?

AI crawler hits on traceten.com by company

Crawl hits by the company that operates the crawler, spoofed requests excluded. 18 crawlers from 11 companies.

  • OpenAIOAI-SearchBot, ChatGPT-User, GPTBot2,18645.5%
  • Googleincludes ordinary Googlebot crawling1,01421.1%
  • Amazon54111.3%
  • Perplexity3326.9%
  • Anthropic2996.2%
  • Meta1873.9%
  • Microsoft1683.5%
  • Apple390.8%
  • DuckDuckGo370.8%
  • Baidu10.0%
  • ByteDance10.0%
Show the data
AI crawler hits on traceten.com by company
CompanyCrawl hitsShare
OpenAI2,18645.5%
Google1,01421.1%
Amazon54111.3%
Perplexity3326.9%
Anthropic2996.2%
Meta1873.9%
Microsoft1683.5%
Apple390.8%
DuckDuckGo370.8%
Baidu10.0%
ByteDance10.0%
Total4,805100.0%

Source: Traceten's AI crawler data for traceten.com, 5 to 29 September 2026 (UTC), queried 30 September 2026. One site, not a benchmark.

OpenAI made 2,186 requests, 45.5% of the total and more than Google, Amazon and Perplexity combined. Most of it was search. OAI-SearchBot, which OpenAI uses to surface sites in ChatGPT's search features, made 1,654 of those requests. ChatGPT-User, which fetches a page when someone asks ChatGPT a question, made 441. GPTBot, the training crawler, made 91.

Google followed with 1,014, then Amazon with 541, Perplexity with 332 and Anthropic with 299. Meta and Microsoft each stayed under 200, Apple and DuckDuckGo under 40, and Baidu and ByteDance appeared once each.

The order of arrival was uneven. Six companies turned up on the first day: Google, OpenAI, Amazon, DuckDuckGo, Apple and Microsoft. Perplexity, Meta and Anthropic followed on the second. Anthropic then went almost silent. After 3 hits on 6 September it made no genuine request until 19 September, when its training crawler, ClaudeBot, fetched robots.txt for the first time. The next day Anthropic made 110 requests.

AI crawler hits per day, 5 to 29 September 2026

All AI crawler hits per UTC day, spoofed requests excluded. The busiest day, 25 September, had 961.

5 Sept: 306 Sept: 1027 Sept: 3418 Sept: 679 Sept: 8310 Sept: 8111 Sept: 7512 Sept: 7513 Sept: 11214 Sept: 11715 Sept: 6616 Sept: 8017 Sept: 17618 Sept: 12119 Sept: 11220 Sept: 23321 Sept: 31022 Sept: 15023 Sept: 15224 Sept: 39125 Sept: 96126 Sept: 22627 Sept: 23428 Sept: 26329 Sept: 247
Show the data
AI crawler hits per day, 5 to 29 September 2026
DateCrawl hits
5 Sept30
6 Sept102
7 Sept341
8 Sept67
9 Sept83
10 Sept81
11 Sept75
12 Sept75
13 Sept112
14 Sept117
15 Sept66
16 Sept80
17 Sept176
18 Sept121
19 Sept112
20 Sept233
21 Sept310
22 Sept150
23 Sept152
24 Sept391
25 Sept961
26 Sept226
27 Sept234
28 Sept263
29 Sept247

Source: Traceten's AI crawler data for traceten.com, 5 to 29 September 2026 (UTC), queried 30 September 2026. One site, not a benchmark.

Volume rose through the month. The first 13 days brought 1,405 hits and the last 12 brought 3,400. The busiest day, 25 September, had 961, of which OpenAI made 494 and Google 346, and almost all of OpenAI's were search crawling. One day is a burst rather than a trend: the four days after it ran between 226 and 263.

Why did they come?

AI crawler hits on traceten.com by purpose

Each crawler has one declared purpose in the Traceten crawler directory. Spoofed requests excluded.

  • Search indexbuilds an AI or web search index2,68255.8%
  • Trainingcollects pages for model training1,04321.7%
  • Other AIno single published purpose59712.4%
  • Answersfetched to answer a question just asked48310.1%
Show the data
AI crawler hits on traceten.com by purpose
PurposeCrawl hitsShare
Search index2,68255.8%
Training1,04321.7%
Other AI59712.4%
Answers48310.1%
Total4,805100.0%

Source: Traceten's AI crawler data for traceten.com, 5 to 29 September 2026 (UTC), queried 30 September 2026. One site, not a benchmark.

Each crawler in the directory has one declared purpose, and the four purposes split very differently:

  • Search index, 2,682 hits (55.8%). Crawlers that build and refresh the search indexes answers draw on, such as OAI-SearchBot, PerplexityBot and Googlebot.
  • Training, 1,043 hits (21.7%). Crawlers that collect pages for model training, such as GPTBot, ClaudeBot and Amazonbot.
  • Other AI, 597 hits (12.4%). Crawlers with no single published purpose, such as GoogleOther and Meta's facebookexternalhit.
  • Answers, 483 hits (10.1%). A page fetched in the moment because a person had just asked an assistant something.

Answer fetches are the smallest group and the most telling, because each one stands for a person's question. ChatGPT-User made 441 of the 483. And 411 of the 483, or 85%, went to the homepage. The 78 blog posts together received 4. Whatever people were asking, the assistant went to the front door. The posts were indexed and read for training, but they were not yet what an assistant fetched to answer a live question.

What did they read?

What AI crawlers requested on traceten.com

Crawl hits by section, spoofed requests excluded. These rows come from the per-page view, which totals 4,803: two short of the daily view, because the two are built separately and one had not caught up when we queried it.

SectionPaths requestedCrawl hitsAnswer fetches
Crawler directory pages1191,37931
Blog posts787664
Homepage1689411
robots.txt156622
Product pages113560
Legal pages63117
Other top-level pages (blog index, directory index, pricing, about)43697
Comparison pages161700
sitemap.xml11500
Everything else, including pages that never existed15461
llms.txt110

After the homepage, the single most requested address was /robots.txt, with 566 requests from nine companies: Google 223, OpenAI 128, Anthropic 89, Meta 46, Microsoft 29, DuckDuckGo 20, Apple 15, Perplexity 15 and ByteDance 1. Google's crawler alone fetched it close to nine times a day. The sitemap was requested 150 times.

/llms.txt was requested once, by Google's search crawler on 10 September. That matches what the state of llms.txt would lead you to expect: a file worth having, and one almost nothing reads.

By section, the crawler directory was read most, with 1,379 hits across 119 pages. The 78 blog posts drew 766, a median of 9 each over 25 days. The six legal pages drew 311, and each of the four most requested ones was read more often than the most requested blog post. A handful of requests asked for pages that have never existed here, among them a coffee brewing guide and a 2024 report.

Were they really who they said?

How each AI crawler hit on traceten.com was verified

The strongest check each hit passed. The 17 spoofed requests are shown apart and are not part of the 4,805.

  • Published IP rangeinside a range the vendor publishes4,53294.3%
  • Operator networkon the company's network, not a crawler list1763.7%
  • User agent onlynothing published to check against962.0%
  • Reverse DNShostname confirmed both ways10.0%
  • Spoofedvendor publishes ranges; address outside them17
Show the data
How each AI crawler hit on traceten.com was verified
Strongest check passedCrawl hitsShare
Published IP range4,53294.3%
Operator network1763.7%
User agent only962.0%
Reverse DNS10.0%
Spoofed (not in the total)17
Total4,805100.0%

Source: Traceten's AI crawler data for traceten.com, 5 to 29 September 2026 (UTC), queried 30 September 2026. One site, not a benchmark.

Of the 4,805 genuine hits, 4,532 came from an address inside a range the crawler's vendor publishes. Another 176 could only be placed on the operator's network. Meta, for example, publishes no crawler ranges, so the best available check is whether the request came from Meta's network. 96 matched a crawler's name with nothing published to check against, and 1 was confirmed by reverse DNS.

The 17 spoofed requests borrowed two companies' names: 12 used OpenAI's and 5 used Anthropic's. By crawler, 11 claimed to be ChatGPT-User, 4 Claude-User, 1 GPTBot and 1 ClaudeBot. Every impostor that claimed to be a user-triggered fetcher asked for a page in our crawler directory, and the two that claimed to be training crawlers asked for the homepage. On 29 September, two requests one second apart claimed to be ChatGPT-User and fetched our GPTBot page. Their user agent differed from the string OpenAI publishes only in where one closing parenthesis sat.

Why this changes how you block

The genuine ChatGPT-User made 441 requests in the same period as the 11 fakes. A rule that blocks on the user agent stops both, and it stops the real one first. Spoofing can also only be detected for vendors that publish ranges. For a crawler matched on its user agent alone, "not spoofed" means nothing contradicted the claim, not that the claim was proven.

Seventeen out of 4,822 requests that used an AI crawler's name is 0.35%. That is small, and it is still not zero on a site almost nobody had heard of.

How do 4,805 crawls compare with 14 visits?

4,805

AI crawler hits, 5 to 29 September

232

Sessions from every source, 30 August to 29 September

14

Sessions referred by an AI assistant in the same window

AI-referred sessions on traceten.com by assistant

Each session is named from its referrer, UTM tags and user agent, with a confidence score and the method that named it.

AssistantSessionsNamed byConfidence
ChatGPT7UTM tag0.99
Claude4UTM tag0.99
Meta AI3User agent0.99

Put side by side, that is about 343 crawler hits for every AI-referred session. For OpenAI alone it is 2,186 hits against 7 ChatGPT sessions. The ratio is striking, and it is easy to read too much into it.

What it does not mean:

  • It is not a conversion rate. A crawler hit is a machine reading a page, and a session is a person arriving. They are different units, and they must never be added together.
  • It is not a delay-free relationship. A page read today may be cited weeks later, and the visit that follows carries no trace of the crawl, which is the crawl-to-citation lag.
  • It is not stable yet. 14 of 232 sessions, or 6.0%, is too few to split by page or to call a trend. At this size, a handful of test clicks from our own team would move it.

What it does mean is that being read is not the same as being recommended. Crawl volume tells you the machines know the site exists. Referrals tell you whether they send people to it. For a young site the gap is wide, and the useful thing is to watch it narrow, month by month, against the same measurement.

What should you do with this?

  1. Keep the search and answer crawlers allowed. OpenAI says sites that opt out of OAI-SearchBot will not be shown in ChatGPT search answers, and Anthropic says disabling Claude-SearchBot or Claude-User may reduce your visibility in its search results. Training crawlers are a separate decision, which the blocking decision table walks through.
  2. Write robots.txt rules per crawler, not per company. OpenAI states that each of its settings is independent, so allowing OAI-SearchBot while disallowing GPTBot is a coherent choice. Anthropic's Claude-SearchBot works the same way.
  3. Verify before you block. The impostors here used exactly the names you would most want to allow. Check the address against what the vendor publishes, then decide.
  4. Measure where the requests arrive. None of this appears in a browser-based analytics tool. AI crawler analytics counts crawls on your server, verifies each one, and keeps spoofed requests apart from the rest.

Frequently asked

01

How often do AI crawlers visit a new website?

On traceten.com, AI crawlers made 4,805 requests in the 25 days from 5 to 29 September 2026, an average of about 192 a day. Volume rose through the month, from 1,405 in the first 13 days to 3,400 in the last 12. That is one young site, so your numbers will depend on your size, links and history.
02

Which AI crawler visits a site most?

On traceten.com it was OpenAI, with 2,186 of 4,805 hits, or 45.5%. Most of those came from OAI-SearchBot (1,654), followed by ChatGPT-User (441) and GPTBot (91). Google was second with 1,014 hits, a total that includes ordinary Googlebot crawling.
03

Do AI crawlers read robots.txt?

Yes, and often. robots.txt was requested 566 times in 25 days by nine companies, more than any page on the site except the homepage. Google's crawler fetched it 223 times, OpenAI's 128 and Anthropic's 89. By comparison, llms.txt was requested once.
04

How common is fake AI crawler traffic?

On traceten.com, 17 of the 4,822 requests that used an AI crawler's name, or 0.35%, came from addresses the named vendor does not publish. They claimed to be OpenAI (12) or Anthropic (5) crawlers. Spoofing can only be detected for vendors that publish IP ranges, so it is a lower bound.
05

Do AI crawler visits turn into website traffic?

Not directly or quickly. traceten.com received 14 AI-referred sessions (7 from ChatGPT, 4 from Claude, 3 from Meta AI) against 4,805 crawler hits in the same period. Crawls and sessions are different units, and a page read today may only be cited, and clicked, weeks later.

Sources and further reading

  1. 01Overview of OpenAI crawlers, OpenAI, checked 30 September 2026
  2. 02Does Anthropic crawl data from the web, and how can site owners block the crawler?, Claude Help Center, checked 30 September 2026
  3. 03AI features and your website, Google Search Central, checked 30 September 2026
Share

Keep reading