Skip to content
Research noteAI-2026-0222

A scoring rule that punishes guessing

Volume-based accuracy rewards a retrieval system for answering everything. A penalty-aware framework separates the three ways it fails.

3 minMain AI Hub

The complaint about retrieval-augmented generation is usually phrased as hallucination, which makes it sound like a model problem. A paper posted to arXiv this week argues it is a measurement problem: a system that answers every question outscores one that declines when its knowledge base cannot support an answer, because the scoring rule counts correct answers and ignores what was said with nothing behind it.

The authors propose scoring that is deliberately asymmetric — a correct answer earns one point, a wrong one loses four, and an abstention earns nothing. Alongside it they plant what they call knowledge-gap canaries: questions whose answers are verifiably absent from the knowledge base, so any answer at all is generation from parametric memory rather than from retrieval. A third component attributes each failure to retrieval, to generation, or to the abstention policy.

Applied to three commercial systems and a no-retrieval baseline on SimpleQA-Verified — a thousand questions, three repeats, graded blind by a three-judge panel drawn from different model families that agreed unanimously 98.9% of the time — the result separates the field on a dimension the usual leaderboard cannot see.

Accuracy when answering was tightly clustered: 97.0% to 98.0% across systems. Canary violation rates differed roughly sixfold, from 16.7% to 98.1%. One system answered almost every question it had no grounds to answer; another mostly declined. On volume-based scoring those two look similar. Under the penalty-aware rule the ranking reorders, and the reordering held across penalty settings from one to nine.

The practical reading for anyone buying such a system is that the accuracy figure in the sales deck is close to meaningless without a companion number for how often the system speaks when it should not. The authors released code, configurations, transcripts and judge votes for independent audit, which is the part that makes the claim checkable rather than merely stated.

Retold from arXiv. This is a summary in our own words; follow the link for the original reporting.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined