Resources / Measurement / Guide
AI influence and incrementality
Assess whether AI answer visibility or a related intervention changed demand beyond a credible comparison, without relabeling unattributed visits as AI traffic.
The short answer
AI answer exposure can affect later demand without creating a measurable referral. A rise in branded search, direct visits, or sales alongside a rise in answer visibility is therefore a reason to investigate, not proof that the answer caused the outcome.
Incrementality asks a narrower question: compared with what would plausibly have happened otherwise, did a defined exposure or intervention change a defined outcome? Answer it with a randomized test when feasible, or with a carefully matched comparison and explicit uncertainty when it is not. This guide teaches how to choose the question, write a study brief, calculate a simple comparison, and recognize when the evidence cannot support a decision. Designing a powered experiment and estimating uncertainty for a real rollout require methods appropriate to its assignment units and data.
Separate recorded referrals from possible influence
An answer-engine referral is a recorded session that matches a source rule. It is useful direct evidence of a visit, subject to missing and misclassified attribution. Website analytics for AI answer-engine journeys explains the cohort and path rules for that evidence.
Possible no-click influence is different. It may be suggested by a stable prompt sample, brand-demand trends, customer research, external panel evidence, or a designed experiment. First-party analytics cannot normally observe that someone saw an answer and later searched for a brand, so it cannot assign that person’s search visit to AI exposure after the fact.
Do not create an “estimated AI traffic” channel by adding direct or branded-search visits to recorded referrals. Preserve the observed channel and report the hypothesis separately.
Choose a study that fits the decision
| Design | When it can help | Main requirement | What the result can support |
|---|---|---|---|
| Randomized exposure or message test | The team controls who sees a comparable AI-related treatment | Random assignment and an outcome measure | A causal estimate for the tested treatment and population |
| Phased rollout | A content, product, or market change must be staged | Predefined rollout order and comparable unexposed groups | A conditional comparison if timing is not driven by expected demand |
| Matched geographic or audience comparison | A randomized rollout is unavailable | A credible untreated comparison with similar pre-period behavior | Evidence consistent with an effect, subject to remaining confounding |
| Interrupted time series | One intervention has a clear date and long stable history | Enough pre- and post-period data and no concurrent disruptive changes | A change in level or trend, with alternative explanations assessed |
| Customer research | The question is discovery memory or self-reported influence | Representative recruitment and neutral wording | Evidence of reported discovery, not verified exposure-to-sale causation |
An answer-visibility prompt sample is a measurement input, not an exposure log. It can show whether the monitored outputs changed under the recorded conditions. It cannot establish who saw those answers.
Decide what treatment you can actually test
Suppose the business asks, “Did our improved buying guide create more qualified inquiries?” A staged page rollout can address that question if the comparison is credible. It cannot isolate AI influence because the page can help visitors from several channels. Changing the label on the report would not change the treatment.
If the narrower question is whether a particular AI-style answer affects consideration, a controlled research test could randomly assign recruited participants to two specified answer presentations and measure a predeclared choice. Keep everything outside the intended treatment comparable, preserve assignment, and record missing responses. Its result concerns those presentations and recruited participants, not natural exposure across answer products or total sales. People who choose to view the answer are not automatically comparable with people who decline; analyzing only willing viewers can undo the purpose of random assignment.
When the team controls neither exposure nor a credible rollout and has only a visibility chart beside a sales chart, start with descriptive monitoring or neutral customer research. That is a useful evidence boundary: a causal estimate is not yet available. Do not fill the gap by relabeling direct traffic.
Pre-register the comparison
Write the following before examining post-change outcomes:
- The intervention or exposure, including dates and who was eligible.
- The primary outcome, such as qualified leads, branded-search clicks, or first purchases.
- The treatment and comparison populations, plus why the comparison is credible.
- The pre-period, post-period, reporting cadence, and expected lag.
- Inclusion, exclusion, identity, and attribution rules.
- Other changes to annotate: campaigns, pricing, launches, seasonality, press, tracking, and competitor activity.
- The minimum decision-relevant effect and the action for a positive, inconclusive, or negative result.
Do not choose the comparison group or cutoff date after seeing the most favorable chart. If the plan changes, version it and treat the altered analysis as exploratory.
Complete a decision brief before the analysis
For the fictional buying-guide study below, a completed brief would specify:
| Choice | Fictional decision and reason |
|---|---|
| Treatment | Improve the warranty comparison on Region A’s buying guide; retain accurate policy access in both regions |
| Comparison | Region B keeps its existing guide during the study; rollout timing is operational, not chosen from expected sales |
| Outcome | Qualified sales inquiries per 10,000 guide-entry sessions, using the same qualification rule in both regions; also show raw counts |
| Window | Equal four-week before/after windows around July 1; review the preceding six weeks for trend differences |
| Main competing explanations | Regional campaigns, price changes, tracking changes, changed traffic composition, and people seeing both versions |
| Decision | Treat +5 inquiries per 10,000 sessions as the minimum worth further investigation, chosen for this fictional decision rather than as a universal benchmark |
| Evidence needed to expand | A credible comparison and an uncertainty assessment that supports a worthwhile effect; a point estimate reaching +5 alone is insufficient |
This brief makes “measure the impact” an inspectable question. If the business instead needs additional inquiry volume, choose that count as the primary outcome and address changes in the eligible population. A per-session rate cannot answer every volume question.
Calculate a difference-in-differences estimate
When a treatment and comparison group had plausibly comparable trends before a staged change, calculate each group’s change, then subtract the comparison change from the treatment change:
incremental change = (treatment after − treatment before) − (comparison after − comparison before)
Worked fictional example
This example uses invented data to demonstrate the comparison arithmetic for a general content intervention. It does not measure AI exposure or establish an AI influence effect. That limitation is the point: a page improvement can affect several channels at once.
A fictional retailer improves a buying guide in Region A on July 1, adding a clear comparison of warranty options and links to the already accurate policy that remains available in both regions. Region B keeps its existing buying guide until the study ends; both regions retain the same accurate warranty facts and policy access. The outcome is qualified sales inquiries per 10,000 buying-guide entry sessions in equal four-week windows. The hypothetical windows each contain 10,000 eligible sessions, so the inquiry counts and normalized rates have the same numerical value. The team had six weeks of similar pre-period direction, no regional campaign changes, and documented that rollout timing was operational rather than demand-led.
| Group | Before: inquiries / sessions | After: inquiries / sessions | Change per 10,000 sessions |
|---|---|---|---|
| Region A, treated | 20 / 10,000 | 29 / 10,000 | +9 |
| Region B, comparison | 18 / 10,000 | 22 / 10,000 | +4 |
First, Region B rises from 18 to 22, a gain of 4 per 10,000 sessions. If that same background change would have applied to Region A without the improvement, A’s expected after value is 20 + 4 = 24. Its observed value is 29, leaving 29 − 24 = 5 above that constructed comparison. This unobserved “what would have happened otherwise” is the counterfactual; Region B helps estimate it rather than revealing it directly.
The equivalent difference-in-differences calculation is (29 − 20) − (22 − 18) = 5 qualified inquiries per 10,000 buying-guide entry sessions. The result is compatible with an effect of the page improvement under the stated comparison. It is not proof that answer engines produced all 5 inquiries: the intervention can help people who arrived through search, email, or direct navigation, and unmeasured regional differences may remain.
The report should also show the raw counts, outcome definition, pre-period trends, traffic denominator, confidence interval or sensitivity range when available, and concurrent changes. If the improvement changes who enters or engages with the buying guide, the per-session denominator can change too; show counts and composition beside the rate. If Region A already had a sharply different upward trend, the comparison is weak and the estimate should be treated as exploratory. Difference-in-differences designs require an explicit parallel-trends assumption; designs with multiple periods or treatment timing need methods suited to those conditions. (Callaway and Sant’Anna’s methodological research)
Test assumptions and competing explanations
Before interpreting the estimate, inspect whether treatment and comparison trends were reasonably parallel before the change, whether tracking or consent changed, whether population composition shifted, and whether major external events differed by group. Re-run the analysis with sensible alternative windows, eligible-population definitions, and exclusion rules. A result that disappears after a defensible sensitivity check should not support a strong decision.
For a randomized study, check assignment, exposure delivery, attrition, and contamination. For a survey, check who was invited, response rate, question wording, recall period, and whether respondents could reasonably distinguish an answer engine from ordinary search.
Recognize a comparison that fails
Imagine a different fictional pre-period: Region A’s weekly rate had already climbed 10 → 15 → 20, while Region B stayed 18 → 18 → 18. That pattern weakens the claim that B represents A’s untreated trend. Subtracting the same +4 background change would not repair it. Investigate the difference and redesign the comparison; do not choose a shorter favorable window after seeing the result.
Similar past trends support a design argument but do not prove that future untreated trends would have stayed parallel. A campaign launched only in A on July 1, or an A-only tracking change, could still explain the apparent effect. A shared buying guide that also changes what B sees can contaminate the comparison.
The worked example supplies arithmetic, not enough evidence to calculate a defensible confidence interval. Two aggregate before/after rows do not establish precision, and 10,000 sessions in each cell do not create 10,000 independently assigned regional treatments. For a real rollout, obtain the repeated observations and assignment-level information a qualified analyst needs. If the effect range remains compatible with no worthwhile gain, report the decision as inconclusive rather than treating a positive point estimate as a win.
Report a bounded conclusion
Use language that matches the design. For the fictional example: “After the staged rollout, Region A increased by an estimated 5 qualified inquiries per 10,000 buying-guide entry sessions more than Region B. The comparison supports a possible incremental effect under the documented assumptions; concurrent unmeasured regional changes remain a limitation.”
Do not say “AI created 5 inquiries” unless a design measured a specific AI exposure and supports that causal claim. Do not use recorded referral revenue as a total influence estimate. Keep direct referral, observed answer visibility, and incremental-study outputs in separate panels.
For this example, the +5 estimate reaches the fictional investigation threshold, but uncertainty and a fully assessed comparison are missing. The defensible next action is to improve or extend the study, not to announce a successful AI campaign. An accurate report can contain a positive estimate and an inconclusive decision.
Decide what to do next
If the result clears the predeclared decision threshold and the assumptions hold, extend or repeat the intervention with a larger or more durable design. If the result is inconclusive, preserve the records and improve the comparison, outcome instrumentation, or observation period. If a factual answer problem motivated the work, preserve it in an answer-capture log and route it through case intake regardless of the demand result.