Resources / Measurement / Guide

AI answer visibility metrics

Calculate bounded presence, citation, recommendation, and representation-accuracy measures from a defined set of observed AI answers.

The short answer

AI answer metrics summarize a defined collection of eligible answer observations. They can show how an entity was represented in that sample under stated conditions. They do not measure all prompts, all users, undisclosed retrieval, website traffic, or incremental business effect.

Start with a versioned prompt set and preserved answer captures. Then define one denominator for each metric, count the matching observations, and report the numerator, denominator, conditions, and exclusions together.

Fictional answer · classification example

1 CedarNote offers shared notebooks. 2 [Source 1]

3 I recommend CedarNote for a team that needs shared online notes.

4 It is the most popular notebook tool among small teams.

  1. Entity mentionThe answer names CedarNote. Confirm the intended entity; classify its role under your rubric.
  2. Visible citationA displayed reference points to a source. Check its destination and which claim it supports.
  3. RecommendationThe answer explicitly endorses an option for a stated need. Naming an option alone would not establish endorsement.
  4. Unsupported claimThe supplied evidence does not establish “most popular.” Record the evidence gap; it does not by itself prove the claim false.

Fictional Source 1: CedarNote’s product specification confirms shared online notebooks. Assume it is an owned-domain page. It supplies no market-share evidence. This is an illustrative source record, not an external citation or a captured product answer.

One answer can contain a mention, an owned-domain citation, a recommendation, and an unsupported claim. Classify each separately; none establishes a referral visit or sale.

Establish the analysis table

Use one row per prompt run, not one row per prompt. Keep raw output in the answer-capture log; the analysis table only carries the classifications needed for calculation.

FieldExample valueWhy it matters
Observation IDOBS-042Lets a reviewer reach the source record
Prompt-set version and stratumNC-1.1, theft coverDefines the sample boundary
Valid answerYesControls eligibility for answer-based rates
Intended entity matchedYesSeparates exact matches from ambiguous names
Qualifying mentionYesApplies the documented meaningful-role rule
Owned-domain citationNoCounts a visible domain citation, not hidden source use
Recommendation outcomeIncluded, not recommendedKeeps inclusion separate from endorsement
Material claim reviewedYes, accurateMakes the accuracy denominator inspectable

Do not turn a refusal, unavailable feature, product error, interrupted capture, or unresolved identity match into a negative answer. Report those outcomes as counts, then state whether the metric excludes them.

Define the denominators before counting

Let E be the number of eligible, valid answers under the named product, conditions, and prompt-set version. Let M be the number of those answers with an exact intended-entity match. Let Q be the number with a qualifying mention under the published rubric. Let C be the number with at least one visible citation to an owned domain. Let CM be the number with both an exact intended-entity match and an owned-domain citation.

Use these measures when their rules match the question:

MetricCalculationMeaning
Run-log completenessplanned runs with a recorded outcome ÷ planned runsWhether every planned attempt has a record, including failures
Eligible-answer completionE ÷ planned runsHow much of the plan supplied answers eligible for the named metric
Entity presence rateM ÷ EThe entity appeared in eligible answers
Qualified mention rateQ ÷ EThe entity appeared in a meaningful defined role
Owned-domain citation rateC ÷ EAn owned domain was visibly cited in an eligible answer
Citation-among-mentions rateCM ÷ MVisible owned citation when the entity was present
Recommendation raterecommended answers ÷ eligible selection-prompt answersExplicit recommendation under a defined selection rule
Claim-accuracy rateaccurate reviewed material claims ÷ reviewed material claimsAccuracy of the claims that were actually reviewed

If a denominator is zero, report the metric as not applicable with its counts, rather than 0%. The citation measures count at most once per answer. If the team needs to count every cited URL or citation marker, publish that as a different event-level metric with its own denominator. Do not call it a citation rate and compare it to an answer-level rate.

Distinguish question coverage from answer presence

Coverage describes the question frame you planned to inspect. In the fictional sampling plan below, 22 selected prompts out of 40 eligible questions give 22 ÷ 40 = 55% question-selection coverage. All four defined strata have selected prompts, giving 4 ÷ 4 = 100% stratum coverage. Neither number says that the entity appeared in an answer or that all real customer questions were represented.

Repeating each selected prompt twice increases the planned observations to 44; it does not increase distinct-question coverage to 110%. Report eligible completions by stratum as well, because planned coverage can hide missing answers in an important group. The sampling guide owns the frame and allocation decisions.

Apply classification rules consistently

An entity is present only after an exact identity match. A name shared by another company, person, or product is ambiguous until review resolves it. A qualifying mention must meet a written role rule, such as a relevant comparison, description, or recommendation. An incidental navigation link or source-list appearance might be a mention but not a qualifying mention.

For claim accuracy, review material claims against the appropriate evidence for that claim. Mark a formerly correct claim as stale when time is the problem. Keep false, incomplete, misleading, and unverifiable as separate outcomes. A visible citation can support, partially support, contradict, or be irrelevant to nearby text; its presence alone does not establish accuracy.

Classify a small batch before calculating a rate

This six-record fictional exercise is separate from the 44-run example below. Its rubric counts an exact entity name in the answer or displayed source label as presence. A qualifying mention must describe, compare, or recommend the entity in the answer itself. A domain citation alone does not establish a name mention. A recommendation requires an explicit favorable selection for the need in a selection prompt.

Record and fictional answerEligible?Present?Qualified?Owned citation?
A: Fact prompt. “Northline Cover covers theft of a locked bicycle subject to the policy’s lock and storage conditions.” Owned policy citation shownYesYesYesYes
B: Selection prompt. “For the locked-bicycle need you described, I recommend Northline Cover; check its policy conditions.” No citationYesYesYesNo
C: Selection prompt. “Compare theft limits and lock requirements.” Source label: “Northline Cover,” linking to its siteYesYesNoYes
D: Selection prompt. “Compare Eastbank’s theft limits before choosing.” Unlabeled northline.example citation shownYesNoNoYes
E: “Service unavailable.” No answerNo———
F: “Northline covers bicycle theft.” No context resolves which NorthlinePending———

Under this predeclared primary rule, E = 4, M = 3, Q = 2, C = 3, and CM = 2. Presence is 3/4 = 75%; owned citation is also 3/4 = 75%, but these are different sets. Citation among mentions is 2/3 = 66.7%, not C/M = 3/3: record D has a citation without a matched name. Only B is recommended among the three eligible selection answers, so recommendation is 1/3 = 33.3%.

This is why reviewers classify records before counting. Two reviewers should apply the rubric to the same preserved answers, explain disagreements, and resolve material cases against the rule. A fresh answer from a new run cannot settle what an older capture said. Keep F pending with its reason, and show the sensitivity to its eventual classification rather than quietly counting it as an absence.

For claim accuracy, make a linked table with one row per material claim: observation ID, exact claim, evidence, and review outcome. Select claims by a stated materiality rule before checking whether they are easy to verify. Several claims from one answer can enter that table; an unreviewed claim must not become “accurate” merely because the answer has a citation.

Calculate a completed fictional example

The following arithmetic uses invented observations from the fictional Northline Cover sampling plan. It is a worked example, not a result about a real insurer or answer product.

The plan called for 44 runs: 22 stable prompts, each captured twice. Forty-one runs returned complete answers: 40 were eligible for the primary exact-entity analysis and 1 had an unresolved entity match. Two other runs ended in product errors and 1 was interrupted. The primary exact-entity rule, defined before collection, excludes the unresolved match from E until review can resolve it. The collection report keeps that record visible and includes a sensitivity result that treats it as a nonmatch; it is not silently discarded as an absence.

Among the 40 eligible answers, 18 had an exact Northline Cover mention, 12 met the qualifying-mention rubric, and 9 visibly cited northline.example. All 9 citation answers also had an exact Northline Cover match, so CM = 9. The team reviewed 25 material claims: 20 accurate, 2 stale, 1 false, 1 incomplete, and 1 unverifiable.

ResultArithmeticReported value
Logged run outcomes44 classified attempts ÷ 44 planned100%; report 2 errors, 1 interrupted, and 1 unresolved match separately
Eligible-answer completion40 ÷ 44 planned90.9%; the four remaining attempts are not eligible answers
Entity presence18 ÷ 4045% of eligible answers
Qualified mention12 ÷ 4030% of eligible answers
Owned-domain citation9 ÷ 4022.5% of eligible answers
Citation among mentionsCM = 9, then 9 ÷ 1850% of entity-present answers
Claim accuracy20 ÷ 2580% of reviewed material claims; outcome mix remains visible

The logged-outcome line shows that every planned attempt received a recorded classification. Eligible-answer completion is 40 of 44, not 100%, because errors, interruption, and unresolved identity do not become eligible answers. The primary answer metrics use 40 eligible answers; the predeclared sensitivity result counts the unresolved answer as a nonmatch: 18 ÷ 41 = 43.9% presence. If it proves to be the intended entity, presence instead becomes 19 ÷ 41 = 46.3%. Both keep the unresolved case visible. The accuracy rate uses 25 reviewed claims, not 40 answers, because one answer can contain several material claims or none. These metrics cannot be averaged into one “visibility score.”

Segment before blending the result

Calculate the same measure by meaningful stratum, product, market, language, or run condition before publishing a total. In the fictional sample, a 45% overall presence rate could conceal 70% presence in comparison prompts and 10% in theft-cover prompts, which need different action.

When the prompt sample deliberately oversamples high-risk groups, a simple overall rate describes the sample allocation, not necessarily the full frame. Weighting can estimate a defined frame only when the frame size, inclusion rules, and weighting method support it. Otherwise label the total “unweighted sample result.” Prompt-set evaluation and sampling explains allocation and frame limits.

Keep peer shares and repeated-run variation inspectable

For a peer comparison, declare the peer set and count each entity at most once per eligible answer under the same qualifying-mention rule. In a separate fictional two-answer batch, answer 1 qualifies Northline and Eastbank; answer 2 qualifies Northline and Southport. Northline has 2 of 4 qualifying entity-answer pairs, or 50% of observed peer mentions. It is present in 2 of 2 answers, or 100% of answers. The two percentages answer different questions; neither is market share. Repeating a name five times in one answer does not add five pairs under this rule.

For repeated-run variation, pair observations of the same prompt under the same conditions. Suppose three prompts have presence outcomes yes/yes, yes/no, and no/no across two runs. One of three pairs disagrees: 33.3% disagreement. Each individual run produces a different presence result, 2/3 and 1/3. Report the pair count and variation instead of presenting the pooled 3/6 as a stable 50% probability. These tiny fictional counts teach the calculation; they do not supply a precision estimate for real monitoring.

Report the result so it can be interpreted

Use a short report statement that preserves the boundary:

In the fictional NC-1.1 stable sample, Northline Cover appeared in 18 of 40 eligible answers (45%) under the recorded product, English-language, signed-out conditions. The result covers the 22 selected prompts and 2 runs each; it does not estimate all customer questions or customer exposure.

Add the date range, product and displayed mode, markets, prompt-set version, eligibility rule, matching rubric, and excluded outcomes beside the table. Keep repeated-run disagreement visible when a result varies. A difference from the previous period is an observed change in the tested sample until a comparison design supports a stronger inference.

Choose the next action from the evidence

An inaccurate material claim should become a case-intake record with the original capture and supporting evidence. A low qualifying-mention rate can trigger content, entity, or source review, but it does not identify the cause by itself. A change in citation rate is not evidence of referral or revenue change; use the website analytics guide for recorded journeys and the influence guide for no-click effects.