Topics / AI answer auditing & correction
AI answer auditing & correction
Auditing turns a questionable answer into a claim that someone can verify. Correction addresses the source or reporting path supported by that evidence, then checks what changed. Start with the observed output: an unfavorable answer is not necessarily inaccurate, and a plausible source is not necessarily its cause.
The work matters when an answer gives people a wrong price, an obsolete qualification, the wrong organization, or another material misunderstanding. Your team can correct information it controls and request review elsewhere. It cannot guarantee what every answer product will say afterward.
Preserve and assess the material issue first
Escalate credible urgent harm immediately. Otherwise, preserve one complete answer, verify the material claim, and classify its consequence before changing sources or submitting feedback. Defer broad monitoring and speculative source changes until the case identifies a supported response. The first task is complete when a responsible owner has an evidence-backed next action and a recheck condition.
Choose the task you have now
| Your situation | Resource | What you should leave with |
|---|---|---|
| You need to preserve the reported answer | Answer-capture log | One complete, immutable observation and its conditions |
| You need to assign and investigate the issue | Case-intake template | A case, responsible owner, evidence gaps, and next action |
| You do not know which issue type or urgency applies | Classification and triage | A supported label and consequence-based response |
| You have verified the problem and need to act | Correction and recheck playbook | Source changes, outside actions, comparable observations, and a closure decision |
| You are reviewing an existing case | Stage-based audit checklist | Recorded results for the relevant case stage |
Keep evidence, findings, actions, and outcomes separate
An effective case separates four records. This prevents a submitted request from becoming an unsupported claim of success.
| Record | What it establishes | What it does not establish |
|---|---|---|
| Observation | The prompt, full answer, visible citations, product, time, and conditions | That the answer is wrong or appears for every user |
| Finding | The disputed claim, verified evidence, classification, severity, and uncertainty | That a suspected source caused the answer |
| Action | A controlled change, outside request, feedback submission, or escalation | That another party accepted it or the answer changed |
| Outcome | The verified source state and later observations under stated conditions | That an action caused the change or resolved every occurrence |
Keep the original observation when the classification changes. Store later answers as new observations. Several runs can support one case; split an answer into separate cases when its claims require different evidence, owners, or responses.
Begin with a case, not a complaint
“The chatbot is wrong” needs an exact passage, affected entity, audience decision, and current evidence before it becomes actionable. Name the person responsible for the next decision. That owner coordinates the work even when a source editor, developer, external publisher, or specialist performs the correction.
A material issue changes a reasonable reader’s decision or understanding. Optional missing detail may not justify correction. Criticism, opinion, ranking, and legitimate expert disagreement require different treatment from a factual error. An organization’s internal record can establish its current offer; it cannot by itself disprove independent criticism.
Preserve the complete answer and relevant source state before editing. Then compare the disputed passage with evidence appropriate to its date, market, and version. Restrict sensitive material to the people who need it. The capture and intake records distinguish the minimum needed to open a case from later investigation fields.
Classify issue type separately from severity
A stale fact was previously correct and is now outdated; use that label only when the historical accuracy is established. A currently contradicted claim with an unknown history can be classified as false with its history marked unknown. An unsupported claim without contrary evidence may instead remain unverifiable. The triage playbook owns the full definitions and decision tree.
Severity depends on credible consequence and urgency. A stale emergency contact could need immediate incident handling, while an obsolete but immaterial description could be low priority. Repetition can affect priority, but one captured answer does not establish audience reach. Preserve a provisional classification while escalating a credible serious risk.
Trace sources without inventing a cause
Inspect visible citations first: do they support the exact claim, entity, and period? Then inspect candidate source layers such as the canonical page, structured data, catalog, feed, maintained profile, and independent coverage.
Matching wording is a reason to investigate, not proof of retrieval. Record whether a source relationship is directly shown, inferred, or unknown. A source can contain an error even when you cannot establish whether that error caused the answer. Fixing a confirmed source defect is still useful work.
Choose the correction you can support
Fix the record that owns a controlled fact and verify its dependent representations. Otherwise a catalog or feed may restore the old value after someone edits the page. For an outside factual error, use the publisher’s documented correction process. For an output problem, use the answer product’s appropriate feedback control. Match formal policy or specialist channels to issues within their scope.
Preserve the submission and any response. “Submitted,” “acknowledged,” and “accepted” describe different states. The correction playbook explains source dependencies, outside routes, sensitive-case handling, and verification with a complete hypothetical case.
Decide what the outcome permits you to say
First check the deployed source; then recheck the answer under comparable product, prompt, market, language, and account conditions. Record material product changes that make comparison imperfect. Set the schedule according to severity and the source’s update path, and store each run separately.
A corrected source and a later changed answer are separate outcomes. If an error still appears in one of three saved runs, the source work may be complete while answer recurrence remains unresolved. If it disappears from all three, the result describes those runs, not every possible answer.
Close when owned actions are complete, outside dependencies and remaining risk have an owner, and a reopen rule exists. Report residual uncertainty explicitly. A closed case can be reopened when evidence changes or a material claim recurs.
Use the case to prevent the next error
A recurring obsolete field is also an ownership or update-process problem. Add the verified lesson to the relevant entity record, content review trigger, technical check, or measurement prompt set. Track work you control, such as verification time and source correction, separately from outside response and answer changes.
Begin with one material claim. Preserve it, verify it, assign a response, and use the correction and recheck playbook to carry the case through to an inspectable result.