Resources / Measurement / Guide

Prompt-set evaluation and sampling guide

Build, test, freeze, and maintain a prompt sample that supports defensible AI-answer measurement without overstating what the observations represent.

The short answer

A useful prompt set represents a defined set of audience questions under recorded conditions. Start with the decision the study must support, define the eligible question population, divide it into meaningful groups, select prompts without favoring convenient or flattering examples, and preserve a stable core for comparison over time.

The result is a sample of observed answers, not a universal score for a brand or model. Report the engines, prompts, markets, dates, run counts, exclusions, and missing observations beside every result.

Use the answer-capture log for each run. Route a material factual problem into the audit case-intake template instead of silently changing its evaluation label.

1. State the measurement question

Write one sentence describing the decision the evaluation should support. Examples include:

Do not combine accuracy, visibility, citations, referrals, and business outcomes into one undefined success score. They have different units, evidence, and limitations.

Then define the unit of analysis. For direct answer review, one unit is usually one prompt run in one named product under one recorded set of conditions. A prompt tested in three products and repeated twice produces six observations, not one.

2. Define the sampling frame

The sampling frame is the documented list or procedure from which prompts can be selected. Build it from evidence of audience demand and organizational risk, such as:

Record the audience, topic, market, language, time period, and source of each candidate. Remove duplicates only after preserving meaningful differences in intent or wording.

A frame can still be incomplete. Private conversations, low-volume questions, new terminology, and prompts used inside products may be unavailable. State those coverage limits instead of calling the frame “all customer questions.”

Turn raw questions into eligible prompt records

Consider these fictional inputs for Northline Cover, the bicycle insurer used in the worked plan below:

Raw inputDecisionReason
“Does Northline cover bike theft?” from a support themeRetain as a broad theft questionA general question is legitimate even when the answer must explain conditions
The identical question copied into a sales spreadsheetMerge the duplicate and retain both source referencesTwo records do not necessarily represent two distinct questions
“Does Northline cover a locked bike stolen from a shared garage?”Retain separatelyThe lock and storage conditions could change the answer
“Will you pay claim 783 for my stolen bike?”Exclude from this public-fact studyAn account-specific decision needs private evidence outside the frame
“Why is Northline the best insurer?” written by the marketing teamRewrite as a neutral comparison or excludeThe wording assumes the preferred conclusion

A retained record might be TC-04 | locked bicycle in shared garage | theft-cover facts | support theme | English | UK | stable-core candidate. Preserve the exact full prompt separately. Ask whether removing or merging a question would erase a condition that changes the answer. If it would, keep the distinction.

A pilot can reveal that a specific-condition test omitted its condition. Add that missing condition and version the item; do not rewrite every broad customer question to make the product look better. If broad and specific needs both belong in the frame, retain both.

3. Divide the frame into meaningful strata

Stratified sampling divides the frame into groups and samples within each group. It helps prevent a large, easy-to-find group from crowding out a small but important one. NIST describes stratification and randomization as ways to reduce systematic sampling error in a sampling scheme. (NIST sampling guidance)

Choose strata that affect the decision. Useful dimensions can include:

DimensionExample strata
IntentLearn, compare, choose, troubleshoot, verify
RelationshipBranded, unbranded category, competitor comparison
TopicProduct, price, policy, people, locations, safety
Journey stageDiscovery, evaluation, purchase, support
RiskRoutine, material, high consequence
MarketCountry, language, region, regulated segment

Do not create every possible combination. Use the smallest set of divisions needed to protect important differences from disappearing in an overall total.

4. Choose the sample and explain its proportions

Use a census when the eligible frame is small enough to test in full. Otherwise, set a target for each stratum before selecting individual prompts.

Three allocation approaches are useful:

Select prompts randomly within a sufficiently large stratum when possible. If a person chooses examples, record the rule and reviewer. A convenience sample made from prompts the team already monitors can support exploration, but it should not be presented as representative.

There is no universal correct prompt count. Precision depends on variability, measurement error, independent repetitions, and the sampling design. Small counts are useful for discovery and directional monitoring, but unstable percentages should be reported with the numerator and denominator, not with false precision.

5. Write prompts without coaching the outcome

Preserve natural wording from the frame when it is available. When prompts must be authored, give the writer the audience need and inclusion rule, not the preferred answer.

For each prompt, record:

Use paraphrase variants only when wording sensitivity is part of the question. Keep the original and each variant as separate prompt IDs linked to one family. Do not average them as though they were independent audience needs.

6. Define evaluation rules before collecting answers

Specify observable fields and decision rules in advance. Typical fields include answer availability, intended-entity presence, mention context, visible citation, cited domain, cited URL, material-claim accuracy, and issue type.

Write a short rubric for every field that requires judgment. Include examples of inclusion, exclusion, ambiguity, and “cannot determine.” Anthropic’s evaluation guidance recommends specific, measurable criteria and task-specific tests that reflect real-world distributions and edge cases. (Anthropic evaluation guidance)

Automated extraction can organize a large collection, but a person should review material facts, ambiguous entity matches, citation support, and classification disagreements. Keep the original answer available so a corrected label does not replace the evidence.

7. Pilot the instrument

Run a small pilot across every stratum before freezing the set. The pilot tests the measurement process, not the brand’s performance.

Check whether:

  1. prompts are understandable without hidden context;
  2. strata and expected facts are applied consistently;
  3. the capture log can preserve the complete output and citations;
  4. two reviewers interpret material labels similarly;
  5. refusals, errors, no-answer results, and ambiguous matches have distinct codes;
  6. the planned run count and review effort are practical;
  7. sensitive or personal information is excluded or protected.

Revise unclear items and document the change. Pilot observations collected under materially different rules should not be merged into the baseline.

Worked fictional sampling plan

This completed example shows the choices a plan needs. Its prompts and results are invented for teaching; they are not research about a real company or answer product.

Decision: Decide whether to open a fact-correction review for a fictional bicycle insurer, Northline Cover, after reports of inaccurate answers about theft cover.

Frame: 40 eligible English-language questions collected from the insurer’s published policy, anonymized support themes, and a customer-question workshop. The frame excludes price quotes, account-specific claims, and questions that require a location the product does not serve.

StratumEligible promptsWhy it is separatePilotStable-core allocation
Theft-cover facts10A false answer can affect a purchase decision26
Exclusions and limits8Omitted limits can mislead25
Comparison questions14Inclusion is a visibility question, not a fact check alone27
Claims and cancellation8Existing customers need operational guidance24
Total40822

The team chose a purposive allocation, giving every stratum at least 4 prompts and additional places to theft, exclusions, and comparison questions. The 6/5/7/4 split reflects that documented priority; it is not a proportional random sample of the whole frame. Report stratum results separately and label the overall result as an unweighted sample summary. To select within a stratum, the reviewer placed its eligible IDs in a spreadsheet, assigned each row a random number, pasted those numbers as fixed values, and sorted the whole table by that column. The first 6 theft IDs, 5 exclusion IDs, 7 comparison IDs, and 4 claims/cancellation IDs became the core. Saving the frozen numbers and selected IDs makes this draw inspectable; recalculating until preferred prompts appear would invalidate the selection rule. Four new policy-language questions were kept in an exploratory set and are reported separately.

The pilot ran each of the 8 prompts once in one named product, signed out, in English, with a UK location setting where available, on the same day. It found that two prompts intended to test locked-bicycle cover did not specify whether the bicycle was locked, so the team rewrote those two prompts to specify the lock condition and versioned the frame from NC-1.0 to NC-1.1. The baseline is 22 stable prompts × 2 separately captured runs = 44 planned observations. Repeated runs show variation under this setup; they do not establish statistical independence or make the 40-prompt frame a census of all audience questions.

Check the plan before collecting the baseline

Try these checks on the fictional plan. The frame has 40 questions, the core has 22, and there are 2 runs per prompt in one product. How many observations are planned? What changes if a second product is added?

The answers are 44 observations, then 88, with 44 reported separately for each product. The number of selected questions remains 22. If a collector replaces a failed run with an easier prompt, the selected sample has changed; retain the failure and follow the predeclared retry rule instead. A retry needs its own record and a stated rule about which attempt enters the primary analysis.

The finished plan should let another collector select the same saved prompt IDs, reproduce the conditions, and explain why each item was included. If it only lists topics and a target prompt count, it is not yet a runnable protocol.

8. Freeze a stable core and keep exploration separate

Version the prompt set before baseline collection. A version record should include the prompt list, sampling frame date, strata, selection method, rubrics, products, conditions, run count, owner, and approval date.

Keep two sets:

Promote an exploratory prompt into a future core version only through documented change control. Do not rewrite a historical prompt because its result is inconvenient. OpenAI’s eval documentation similarly treats success criteria, test data, and runs as explicit parts of an evaluation and uses representative test data for prompt testing. (OpenAI evals guide)

9. Run under recorded conditions

Define the product, mode, market, language, date range, login state, subscription tier, personalization state when known, device, location setting, conversation state, and whether search or another tool is active.

Randomize run order where order or time could create a systematic difference. Start a new conversation when independence from earlier turns is required. If a product does not reveal a condition, mark it unknown.

Repeat a controlled subset when output variability matters. Store each run separately. Repetition estimates variability within the tested conditions; it does not turn a narrow prompt frame into a representative market sample.

Separate reasoning settings and conversation journeys

Record the selected reasoning or effort setting when the interface exposes it; otherwise record “unknown.” Keep changed settings in separate comparison groups. Do not infer an unseen mode from the length or tone of an answer.

For multi-turn evaluation, define the starting prompt, follow-up sequence, branching rule, and stopping point before collection. Preserve every turn and the preceding conversation. A follow-up that names your brand measures behavior after that cue; it is not an unprompted discovery observation.

Fictional protocol: Run ten independent conversations per product, each beginning “How should a small team choose a shared notebook?” Then ask “Which options support offline work?” without inserting a brand. Record mention and citation presence separately for each turn. If a brand appears in six first turns and four second turns, report 6/10 and 4/10 for the respective stages. Those are paired stages from ten journeys, not twenty independent users. To calculate retention, count how many of the original six also appear in the second turn; the two totals alone cannot establish that overlap.

Set cadence from the decision and observed variability

Pilot repeated runs within a collection window and repeat comparable windows before choosing an ongoing cadence. Separate variability between runs in the same window from changes between windows. Review sooner when product changes, material errors, or an approaching decision make a delay costly. A provider-wide source-churn statistic cannot prescribe a schedule for your own prompt panel.

Record the cadence, rationale, budget, and review trigger per product. If cadence changes, document the break and compare equivalent windows rather than pooling unequal observation opportunities. These are study-design recommendations, not provider requirements or a promise that repeated observations capture every user’s experience.

10. Analyze with the correct denominator

For every rate, name the numerator and denominator. Examples include:

Do not quietly remove refusals, timeouts, product errors, unavailable features, or ambiguous matches. Report them as separate outcomes and explain which measures exclude them.

Break results out by meaningful stratum before presenting a total. A stable aggregate can hide a severe decline in one market or a small high-risk group. When the design intentionally oversamples a stratum, either weight estimates back to the frame or label totals as unweighted sample results.

Use the AI answer visibility metrics guide to calculate the resulting presence, citation, recommendation, and claim-accuracy rates. That guide owns metric math; this page owns the sampling design.

Copyable sampling plan

Study name and version:
Decision this study supports:
Measurement question:
Unit of analysis:

Audience and scope:
Markets and languages:
Answer products and modes:
Observation window:

Sampling frame sources:
Frame inclusion rules:
Frame exclusions and known gaps:

Strata and rationale:
Allocation method:
Target prompts by stratum:
Selection method within each stratum:
Stable-core prompt IDs:
Exploratory prompt IDs:

Runs per prompt:
Recorded conditions:
Reasoning/effort setting (or unknown):
Conversation protocol, turns, branching, and stopping rule:
Run-order rule:
Missing-observation codes:

Fields and rubrics:
Material facts and approved evidence:
Manual-review rule:
Disagreement-resolution rule:

Pilot date and findings:
Baseline date:
Cadence by product and rationale:
Cadence review trigger:
Change-control owner:
Retention period:

Known limitations:
Approved by and date:

Validation checklist

Before reporting, confirm that: