Resources / Measurement / Guide
AI answer visibility metrics
Calculate bounded presence, citation, recommendation, and representation-accuracy measures from a defined set of observed AI answers.
The short answer
AI answer metrics summarize a defined collection of eligible answer observations. They can show how an entity was represented in that sample under stated conditions. They do not measure all prompts, all users, undisclosed retrieval, website traffic, or incremental business effect.
Start with a versioned prompt set and preserved answer captures. Then define one denominator for each metric, count the matching observations, and report the numerator, denominator, conditions, and exclusions together.
1 CedarNote offers shared notebooks. 2 [Source 1]
3 I recommend CedarNote for a team that needs shared online notes.
4 It is the most popular notebook tool among small teams.
- Entity mentionThe answer names CedarNote. Confirm the intended entity; classify its role under your rubric.
- Visible citationA displayed reference points to a source. Check its destination and which claim it supports.
- RecommendationThe answer explicitly endorses an option for a stated need. Naming an option alone would not establish endorsement.
- Unsupported claimThe supplied evidence does not establish “most popular.” Record the evidence gap; it does not by itself prove the claim false.
Fictional Source 1: CedarNote’s product specification confirms shared online notebooks. Assume it is an owned-domain page. It supplies no market-share evidence. This is an illustrative source record, not an external citation or a captured product answer.
Establish the analysis table
Use one row per prompt run, not one row per prompt. Keep raw output in the answer-capture log; the analysis table only carries the classifications needed for calculation.
| Field | Example value | Why it matters |
|---|---|---|
| Observation ID | OBS-042 | Lets a reviewer reach the source record |
| Prompt-set version and stratum | NC-1.1, theft cover | Defines the sample boundary |
| Valid answer | Yes | Controls eligibility for answer-based rates |
| Intended entity matched | Yes | Separates exact matches from ambiguous names |
| Qualifying mention | Yes | Applies the documented meaningful-role rule |
| Owned-domain citation | No | Counts a visible domain citation, not hidden source use |
| Recommendation outcome | Included, not recommended | Keeps inclusion separate from endorsement |
| Material claim reviewed | Yes, accurate | Makes the accuracy denominator inspectable |
Do not turn a refusal, unavailable feature, product error, interrupted capture, or unresolved identity match into a negative answer. Report those outcomes as counts, then state whether the metric excludes them.
Define the denominators before counting
Let E be the number of eligible, valid answers under the named product, conditions, and prompt-set version. Let M be the number of those answers with an exact intended-entity match. Let Q be the number with a qualifying mention under the published rubric. Let C be the number with at least one visible citation to an owned domain. Let CM be the number with both an exact intended-entity match and an owned-domain citation.
Use these measures when their rules match the question:
| Metric | Calculation | Meaning |
|---|---|---|
| Run-log completeness | planned runs with a recorded outcome ÷ planned runs | Whether every planned attempt has a record, including failures |
| Eligible-answer completion | E ÷ planned runs | How much of the plan supplied answers eligible for the named metric |
| Entity presence rate | M ÷ E | The entity appeared in eligible answers |
| Qualified mention rate | Q ÷ E | The entity appeared in a meaningful defined role |
| Owned-domain citation rate | C ÷ E | An owned domain was visibly cited in an eligible answer |
| Citation-among-mentions rate | CM ÷ M | Visible owned citation when the entity was present |
| Recommendation rate | recommended answers ÷ eligible selection-prompt answers | Explicit recommendation under a defined selection rule |
| Claim-accuracy rate | accurate reviewed material claims ÷ reviewed material claims | Accuracy of the claims that were actually reviewed |
If a denominator is zero, report the metric as not applicable with its counts, rather than 0%. The citation measures count at most once per answer. If the team needs to count every cited URL or citation marker, publish that as a different event-level metric with its own denominator. Do not call it a citation rate and compare it to an answer-level rate.
Distinguish question coverage from answer presence
Coverage describes the question frame you planned to inspect. In the fictional sampling plan below, 22 selected prompts out of 40 eligible questions give 22 ÷ 40 = 55% question-selection coverage. All four defined strata have selected prompts, giving 4 ÷ 4 = 100% stratum coverage. Neither number says that the entity appeared in an answer or that all real customer questions were represented.
Repeating each selected prompt twice increases the planned observations to 44; it does not increase distinct-question coverage to 110%. Report eligible completions by stratum as well, because planned coverage can hide missing answers in an important group. The sampling guide owns the frame and allocation decisions.
Apply classification rules consistently
An entity is present only after an exact identity match. A name shared by another company, person, or product is ambiguous until review resolves it. A qualifying mention must meet a written role rule, such as a relevant comparison, description, or recommendation. An incidental navigation link or source-list appearance might be a mention but not a qualifying mention.
For claim accuracy, review material claims against the appropriate evidence for that claim. Mark a formerly correct claim as stale when time is the problem. Keep false, incomplete, misleading, and unverifiable as separate outcomes. A visible citation can support, partially support, contradict, or be irrelevant to nearby text; its presence alone does not establish accuracy.
Classify a small batch before calculating a rate
This six-record fictional exercise is separate from the 44-run example below. Its rubric counts an exact entity name in the answer or displayed source label as presence. A qualifying mention must describe, compare, or recommend the entity in the answer itself. A domain citation alone does not establish a name mention. A recommendation requires an explicit favorable selection for the need in a selection prompt.
| Record and fictional answer | Eligible? | Present? | Qualified? | Owned citation? |
|---|---|---|---|---|
| A: Fact prompt. “Northline Cover covers theft of a locked bicycle subject to the policy’s lock and storage conditions.” Owned policy citation shown | Yes | Yes | Yes | Yes |
| B: Selection prompt. “For the locked-bicycle need you described, I recommend Northline Cover; check its policy conditions.” No citation | Yes | Yes | Yes | No |
| C: Selection prompt. “Compare theft limits and lock requirements.” Source label: “Northline Cover,” linking to its site | Yes | Yes | No | Yes |
D: Selection prompt. “Compare Eastbank’s theft limits before choosing.” Unlabeled northline.example citation shown | Yes | No | No | Yes |
| E: “Service unavailable.” No answer | No | — | — | — |
| F: “Northline covers bicycle theft.” No context resolves which Northline | Pending | — | — | — |
Under this predeclared primary rule, E = 4, M = 3, Q = 2, C = 3, and CM = 2. Presence is 3/4 = 75%; owned citation is also 3/4 = 75%, but these are different sets. Citation among mentions is 2/3 = 66.7%, not C/M = 3/3: record D has a citation without a matched name. Only B is recommended among the three eligible selection answers, so recommendation is 1/3 = 33.3%.
This is why reviewers classify records before counting. Two reviewers should apply the rubric to the same preserved answers, explain disagreements, and resolve material cases against the rule. A fresh answer from a new run cannot settle what an older capture said. Keep F pending with its reason, and show the sensitivity to its eventual classification rather than quietly counting it as an absence.
For claim accuracy, make a linked table with one row per material claim: observation ID, exact claim, evidence, and review outcome. Select claims by a stated materiality rule before checking whether they are easy to verify. Several claims from one answer can enter that table; an unreviewed claim must not become “accurate” merely because the answer has a citation.
Calculate a completed fictional example
The following arithmetic uses invented observations from the fictional Northline Cover sampling plan. It is a worked example, not a result about a real insurer or answer product.
The plan called for 44 runs: 22 stable prompts, each captured twice. Forty-one runs returned complete answers: 40 were eligible for the primary exact-entity analysis and 1 had an unresolved entity match. Two other runs ended in product errors and 1 was interrupted. The primary exact-entity rule, defined before collection, excludes the unresolved match from E until review can resolve it. The collection report keeps that record visible and includes a sensitivity result that treats it as a nonmatch; it is not silently discarded as an absence.
Among the 40 eligible answers, 18 had an exact Northline Cover mention, 12 met the qualifying-mention rubric, and 9 visibly cited northline.example. All 9 citation answers also had an exact Northline Cover match, so CM = 9. The team reviewed 25 material claims: 20 accurate, 2 stale, 1 false, 1 incomplete, and 1 unverifiable.
| Result | Arithmetic | Reported value |
|---|---|---|
| Logged run outcomes | 44 classified attempts ÷ 44 planned | 100%; report 2 errors, 1 interrupted, and 1 unresolved match separately |
| Eligible-answer completion | 40 ÷ 44 planned | 90.9%; the four remaining attempts are not eligible answers |
| Entity presence | 18 ÷ 40 | 45% of eligible answers |
| Qualified mention | 12 ÷ 40 | 30% of eligible answers |
| Owned-domain citation | 9 ÷ 40 | 22.5% of eligible answers |
| Citation among mentions | CM = 9, then 9 ÷ 18 | 50% of entity-present answers |
| Claim accuracy | 20 ÷ 25 | 80% of reviewed material claims; outcome mix remains visible |
The logged-outcome line shows that every planned attempt received a recorded classification. Eligible-answer completion is 40 of 44, not 100%, because errors, interruption, and unresolved identity do not become eligible answers. The primary answer metrics use 40 eligible answers; the predeclared sensitivity result counts the unresolved answer as a nonmatch: 18 ÷ 41 = 43.9% presence. If it proves to be the intended entity, presence instead becomes 19 ÷ 41 = 46.3%. Both keep the unresolved case visible. The accuracy rate uses 25 reviewed claims, not 40 answers, because one answer can contain several material claims or none. These metrics cannot be averaged into one “visibility score.”
Segment before blending the result
Calculate the same measure by meaningful stratum, product, market, language, or run condition before publishing a total. In the fictional sample, a 45% overall presence rate could conceal 70% presence in comparison prompts and 10% in theft-cover prompts, which need different action.
When the prompt sample deliberately oversamples high-risk groups, a simple overall rate describes the sample allocation, not necessarily the full frame. Weighting can estimate a defined frame only when the frame size, inclusion rules, and weighting method support it. Otherwise label the total “unweighted sample result.” Prompt-set evaluation and sampling explains allocation and frame limits.
Keep peer shares and repeated-run variation inspectable
For a peer comparison, declare the peer set and count each entity at most once per eligible answer under the same qualifying-mention rule. In a separate fictional two-answer batch, answer 1 qualifies Northline and Eastbank; answer 2 qualifies Northline and Southport. Northline has 2 of 4 qualifying entity-answer pairs, or 50% of observed peer mentions. It is present in 2 of 2 answers, or 100% of answers. The two percentages answer different questions; neither is market share. Repeating a name five times in one answer does not add five pairs under this rule.
For repeated-run variation, pair observations of the same prompt under the same conditions. Suppose three prompts have presence outcomes yes/yes, yes/no, and no/no across two runs. One of three pairs disagrees: 33.3% disagreement. Each individual run produces a different presence result, 2/3 and 1/3. Report the pair count and variation instead of presenting the pooled 3/6 as a stable 50% probability. These tiny fictional counts teach the calculation; they do not supply a precision estimate for real monitoring.
Report the result so it can be interpreted
Use a short report statement that preserves the boundary:
In the fictional
NC-1.1stable sample, Northline Cover appeared in 18 of 40 eligible answers (45%) under the recorded product, English-language, signed-out conditions. The result covers the 22 selected prompts and 2 runs each; it does not estimate all customer questions or customer exposure.
Add the date range, product and displayed mode, markets, prompt-set version, eligibility rule, matching rubric, and excluded outcomes beside the table. Keep repeated-run disagreement visible when a result varies. A difference from the previous period is an observed change in the tested sample until a comparison design supports a stronger inference.
Choose the next action from the evidence
An inaccurate material claim should become a case-intake record with the original capture and supporting evidence. A low qualifying-mention rate can trigger content, entity, or source review, but it does not identify the cause by itself. A change in citation rate is not evidence of referral or revenue change; use the website analytics guide for recorded journeys and the influence guide for no-click effects.