AI abstention design
Adjust abstention thresholds with observed task outcomes
Changing a threshold alters which questions receive answers. Review newly accepted and newly withheld cases before applying the change broadly.
In this article
Preserve the current decision rule
Record the existing threshold or rule set with its task scope, source configuration and evaluation results. A score threshold is meaningful only in the context of the system producing the score.
Identify why the change is needed. Excessive refusal, unsupported answers and poor clarification behaviour may require different fixes. Lowering one threshold is not a universal remedy.
Keep authority and required-approval rules deterministic. They should not be relaxed as part of a confidence experiment.
Compare the boundary cases
Run the candidate rule on the same reviewed dataset and inspect cases whose outcomes change. Newly answered cases deserve particular attention because they were previously withheld for a reason.
Check whether the evidence supports the full answer, including conditions and exceptions. A relevant passage is not necessarily sufficient for the requested conclusion.
For newly withheld cases, determine whether clarification or a limited answer could preserve useful work. A stricter binary refusal can be easier to implement while producing a worse experience.
Roll out by a meaningful task class
Enable the change for a defined category and observe correct completion, unsupported answers, unnecessary refusal and escalation workload. Keep the configuration version with each result.
Use reviewed samples from the actual cohort, including apparently successful answers. Relying only on complaints misses errors users do not recognise.
Provide a way to restore the previous decision rule without changing unrelated prompts or retrieval settings. This helps isolate the effect and makes restoration easier to evaluate.
Revisit the underlying signal
If a threshold behaves inconsistently across categories, inspect whether the score measures the property the product needs. A retrieval score may identify topical similarity while failing to establish answer sufficiency.
Consider explicit checks for required evidence or conflicting versions rather than endlessly tuning a broad confidence number. Keep the rule understandable enough that support can explain why a task was withheld.
Expand only after the observed tradeoff meets the intended use. The goal is appropriate answers and useful next steps, with a clear record of which cases the changed rule now accepts or declines.
Primary sources
Microsoft Learn: retrieval and answer evaluatorsMicrosoft Learn: evaluation and observabilityReferences checked 11 September 2026.