Grounded answer citations

Score claim support separately from polished writing

A fluent answer can still misstate its evidence. Measure whether its claims are supported, whether important conditions survived and whether readers can inspect the source.

In this article

Choose the unit you will review

An answer-level score is convenient but coarse. One paragraph might contain four correct statements and one unsupported instruction that changes what a reader does. For consequential questions, break the answer into claims and review each against the evidence supplied to the model.

Define what counts as a claim before collecting results. Dates, amounts, eligibility rules and recommendations based on source material are obvious candidates. Transitional phrases generally are not. Reviewers should not produce different denominators simply because one splits a sentence more aggressively than another.

Keep the question and expected essential facts with the review. A system can achieve high support by saying almost nothing. Supported claims alone do not measure whether the answer addressed the user's need.

Track distinct failure types

Reference validity asks whether the citation points to evidence actually supplied. Claim support asks whether that evidence justifies the statement. Coverage asks whether the answer includes the facts needed to answer the question. Source accessibility asks whether the intended reader can inspect it. Report these separately.

For each unsupported claim, record a reason. Useful categories include contradiction, missing qualification, unsupported inference and evidence too weak to decide. The remedies differ. A missing exception may require better context selection, while an invented percentage may require generation and validation changes.

Do not treat a reachable URL as evidence of factual support. Link health belongs in operational checks. It is valuable, but it cannot replace reading the cited passage.

Use a reviewed sample to calibrate automation

Begin with human review of representative questions. Include ordinary questions and deliberately difficult ones involving conflicting revisions, exceptions and incomplete sources. Mark disagreement explicitly and resolve unclear review instructions before comparing model versions.

If an automated groundedness evaluator is used, measure its agreement with the reviewed sample. Inspect false passes especially carefully. A high average agreement can coexist with poor performance on the specific exception cases that matter to the business.

Publish sample sizes with rates. Ten reviewed answers and a thousand reviewed answers do not provide the same confidence. If rare but serious failures are present, report the examples and their consequences rather than allowing them to disappear in an aggregate percentage.

Connect the result to a release decision

Set acceptance criteria around the intended use. A discovery assistant may tolerate an explicit incomplete answer. A system supporting an operational decision may require every actionable claim to be supported or referred for review. Decide this before seeing the candidate's score.

Compare releases on the same questions, source snapshot and review rules. Also measure refusal and unanswered rates so a candidate cannot appear safer simply by avoiding useful answers. The release record should explain which failures remain, how the interface handles them and who accepts that behaviour. A single quality number cannot carry those decisions on its own.

Primary sources

Microsoft Learn: groundedness evaluation

References checked 11 September 2026.