A funnel report is an average. It takes every visitor who reached step one, follows them down, and tells you what fraction survived each stage. That is useful right up to the moment two groups inside it stop behaving alike, and then the average describes nobody, moves for reasons you cannot name, and quietly hides the thing you were trying to find.
Visitors an AI assistant sent you are one of those groups. They arrive having already had a conversation about your category with something that summarised four vendors and recommended one. They skipped the comparison stage you built. They land on the page that got cited rather than the page you optimised. Whether that makes them better or worse is an empirical question, and there is exactly one honest way to answer it: run the funnel twice, once for them and once for everyone else, and look at where the two lines separate.
The question an aggregate funnel cannot answer
Standard analytics gives you one funnel and a source report sitting next to it. The source report says 6% of your sessions came from AI assistants. The funnel says 3.1% of visitors who hit pricing eventually signed up. Both are true. Neither tells you whether those 6% signed up at 1% or at 11%, and no amount of staring at the two charts together will produce the answer, because the join that would produce it does not exist in either one.
What you need is the funnel computed within the segment. Same steps, same window, different population. Then the comparison is a like-for-like one and the difference between the two curves is attributable to the thing you varied.
Traceten reports the top sources for every step of a funnel, up to eight of them, which is the same computation viewed from the other end: instead of filtering the funnel down to one source, it shows you the source mix at each step and lets you watch that mix change as the funnel narrows. A source that is 14% of step one and 4% of step four is being lost somewhere in between, and the step where its share collapses is the step that failed it.
Where the two readings differ
Both use the same underlying step records. The aggregate answers 'how many'; the segmented one answers 'which of them, and where'.
Aggregate funnel
One curve. Tells you the overall rate and the biggest absolute drop. Cannot tell you which population produced either.
Per-step source mix
The share each source holds at each step. A share that falls as you go down names the step where that source is losing people.
Segmented funnel
The same funnel run over one source's visitors. Gives you a rate you can compare directly against the site baseline.
The three shapes, and what each one means
Once you have two curves on the same axes, almost everything you will see is one of three shapes. They demand completely different responses, and the aggregate funnel renders all three identically.
Parallel, offset. The AI-referred curve sits above or below the baseline at every step and decays at the same rate. There is nothing wrong with any particular step. The difference happened before the funnel started: these visitors arrived with a different level of intent, and that level carried through unchanged. The action here is not a funnel fix. If they are converting better, the finding is that the channel is under-invested. If worse, the finding is that the assistant is sending people who were never going to buy, and the fix lives in what your pages say about who the product is for.
Divergence at one step. The curves track each other and then split at a single stage. This is the finding worth having, because it names a specific screen. The usual cause is context: the assistant told the visitor something your page then contradicts, or assumed something your page requires them to already know. A visitor who was told "it integrates with Shopify" and lands on a page that does not mention Shopify anywhere drops at exactly the step where they went looking for confirmation.
Crossover. The AI segment starts ahead and finishes behind. They engage more and buy less. This is the expectation-gap signature, and it is the one worth acting on fastest, because every one of those visitors is a person the assistant recommended you to and who then decided against you. Whatever they went looking for and did not find is a sentence missing from a page you own.
Reading the shape
The aggregate funnel produces the same summary number for all three. The segmented view is what makes them distinguishable.
| Shape | What it means | Where the fix lives | What not to do |
|---|---|---|---|
| Parallel, AI curve above | Better-qualified arrivals, no step-level problem | Budget and content that earns more of these arrivals | Rewriting a step that is working fine for everyone |
| Parallel, AI curve below | The assistant is recommending you to the wrong people | The pages that describe who the product is for | Treating it as a landing-page conversion problem |
| Divergence at one step | Something at that step contradicts what the visitor was told | That specific page, and only that one | A site-wide redesign off a single-step signal |
| Crossover | Expectation gap: high engagement, low completion | The claim the assistant is making that you do not substantiate | Blaming the traffic quality before reading the answers |
The mistake that makes the whole exercise impossible
Here is the failure that costs teams a quarter, and it happens before any data is collected.
The funnel's cohort is fixed by its first step. The visitors in a funnel are the ones whose first step falls inside your date range; every later step is counted for those people and nobody else. That rule is what makes a funnel a funnel rather than four unrelated counts, but it means the first step is not a neutral choice. It is the population definition.
So a funnel that begins at the homepage measures people who visited the homepage. And AI-referred visitors, disproportionately, do not. They arrive on the page the assistant cited: a comparison page, a docs page, a specific feature page, a blog post that answered the question they asked. That is the entire mechanic of how AI referrals reach you: the citation is a deep link, not a front-door referral.
Build the funnel homepage-first and the AI segment shows up as a rounding error, not because the channel is small but because you defined the cohort in a way that excludes it. Then somebody concludes AI traffic does not convert, from a funnel that structurally could not contain it.
Check this before you read a single number
Pull your top landing pages for AI-referred sessions. If the first step of your funnel is not one
of them, and does not contains-match one of them, your funnel is not measuring this channel.
Change the first step before you change anything else.
The fix is to start the funnel where the traffic actually lands. A starts_with match on /docs or a contains match on /compare will hold a cohort that a homepage step never could. You will often end up with two funnels (a front-door one and a deep-landing one), and that is the correct outcome rather than a compromise, because they genuinely are two different journeys and averaging them was the original problem.
The window is part of the definition
Every funnel has a conversion window: the time a visitor has to get from the first step to the last. Seven days by default, adjustable from one hour to ninety.
The window is not a display setting. It is part of what the number means, and it interacts with the segmentation in a way that catches people out. Consideration length is a property of the population, not of the funnel, and if AI-referred visitors are arriving further along in their decision, they may complete in a window that is genuinely shorter than your baseline. Run both segments at a seven-day window and the comparison is fair. Run one at seven days and the other at thirty because somebody adjusted it mid-analysis, and the difference you are looking at is the window.
There is a second interaction worth knowing. Visitors are recruited by a first step inside your date range, and their later steps are counted for up to one window past the end of it. So a seven-day funnel over a seven-day range is mostly measuring people whose window has not closed. Widen the range or shorten the window; do not read the last step.
Small numbers, and the discipline they require
Segmenting cuts your sample. That is the cost of the method and there is no way around it.
If AI referrals are 6% of your traffic and your funnel's first step catches 3,000 visitors, the AI segment is about 180 people at step one and perhaps 20 by step four. A step-four conversion rate computed on 20 visitors moves by five percentage points when one person changes their mind.
Two rules keep this honest, and both are unglamorous:
Widen the range, not the claim. A 180-day results range is available; use it before you conclude anything about a step whose segment count is in the tens. A quarter of data on a small segment is a better instrument than a week of data and a confident sentence.
Report the count next to the rate, always. "AI-referred visitors convert at 8.2%" is a claim that cannot be evaluated. "18 of 220, so 8.2%" is a claim a reader can weigh, and the second version is the one that survives somebody senior asking how sure you are. The same discipline governs every rate we publish: a percentage without its denominator is decoration.
2 to 8
Steps per funnel. Fewer steps hold a larger segment at the bottom, which is the constraint that actually binds when you segment
start shallow, deepen later
180 days
Maximum results range. Reach for it early when the segment is small, rather than defending a number computed on twenty people
range is free, confidence is not
8
Sources reported per step. Enough to watch a mix shift; not enough to make a long-tail source visible, so check the source report too
top-8, not all
Countries, and why the list does not change per step
One detail that looks like a quirk and is actually a privacy property worth understanding, because it will shape how you read the geographic breakdown.
Each step reports its top countries. Countries below a minimum visitor count are not named. They are folded into an other bucket rather than dropped, so the totals still reconcile. And the set of named countries is decided once for the whole funnel, not per step.
That second part is the load-bearing one. Step counts only ever fall as you go down a funnel. If a country were named at step one and omitted at step four for being too small, its step-four count would be recoverable by subtraction from the other bucket, which is precisely the small-number disclosure the threshold exists to prevent. Deciding the named set once, for the whole funnel, closes that. It is the same reasoning that governs attribution without an identity graph: the useful version of a privacy rule is the one that survives someone doing arithmetic on the output.
What to actually do this week
Pick your one real conversion. Build a three-step funnel whose first step matches where AI traffic lands, not where you wish it landed. Set the window to something shorter than your date range. Read the per-step source mix rather than the headline rate, and look for the step where the AI share collapses.
If nothing collapses and the curves are parallel, you have learned something real and you are done: the channel differs in level, not in shape, and the work is upstream in what gets you cited rather than downstream in the funnel.
If something does collapse, you have a page and a hypothesis, and the hypothesis is nearly always the same one: the visitor was told something, and that page did not confirm it.
Frequently asked
Frequently, but the direction is site-specific and the useful finding is rarely the rate. Because AI-referred visitors arrive after a conversation that already compared vendors, they tend to enter mid-consideration and land on a deep page rather than a homepage. That changes which step they drop at more reliably than it changes the overall percentage. Run the same funnel over the AI segment and the site baseline, and read where the two curves separate rather than comparing two summary numbers.
Almost always because the funnel's first step is your homepage. A funnel's cohort is fixed by its first step, so it can only contain visitors who took that step, and AI-referred visitors usually land on the specific page the assistant cited instead. Pull your top landing pages for AI-referred sessions and make the first step match one of them, using a starts_with or contains operator rather than an exact path.
Two ways, and they answer slightly different questions. Read the per-step source breakdown, which shows the share each source holds at every step and makes a collapsing share visible directly. Or run the funnel over a single source's visitors and compare the resulting curve against the site baseline, which gives you a rate you can quote. Use the first to find the step, the second to size the effect.
Shorter than your date range, and identical across every segment you intend to compare. Visitors are recruited by a first step inside the range and their later steps are counted for up to one window beyond it, so a seven-day window over a seven-day range leaves most visitors still inside their window and the final step reads artificially low. Default to seven days, and change it only for a reason you can state.
There is no universal threshold, but a rate computed on fewer than about 30 visitors at a step will move by several percentage points on one person's decision, which makes it unusable for comparing two segments. Widen the date range before narrowing the conclusion, and always publish the count beside the rate so a reader can judge it themselves.
Sources & further reading
- 01Referrer-Policy, MDN Web Docs
- 02AI Insights: crawl and referral traffic, Cloudflare Radar
- 03GEO: Generative Engine Optimization, Aggarwal et al., KDD 2024 (arXiv:2311.09735)
- 04Simpson's paradox, Stanford Encyclopedia of Philosophy