A benchmark measuring how often retrieval-grounded models answer confidently when no supporting evidence exists, plus a primitive that turns silent confabulation into an auditable refusal signal.
A benchmark measuring how often retrieval-grounded models answer confidently when no supporting evidence exists, plus a primitive that turns silent confabulation into an auditable refusal signal.
Project Details
Updated 07/13/26 · Edited by orgRetrieval-grounded language models fail silently when the evidence isn't there. Asked something that the underlying sources don't support, they will interpolate a fluent confident answer from priors rather than actively reporting that nothing supports the claim. This is a metric standard accuracy metrics can't catch because the confabulated answer looks exactly like a grounded one. This project builds a benchmark that measures that failure directly through a set of scenarios where the correct behavior is to report absence, scored on how often a system is the idea produces a confident claim over an empty evidence set. The approach comes out of the infrastructure we've built at QuarterMill where absence is a typed, first-class object rather than something inferred from silence, and where claims resolve into sources in custody.
The work is led by Keith G. Pemberton II and a team of undergraduate and graduate researchers with experiences at Baidu, Perplexity, and Duolingo. The output is a public benchmark, an open eval harness, and a write up of what current grounded-generation systems do when the evidence is empty.
Theory of Impact
Updated 07/21/26 · By grantmaking.aiAs AI Systems are promoted to higher stakes institutional decisions, more of them are retrieval grounded, yet the safety story rests on the model being tied to real evidence. In reality, grounding fails silently. A model that confabulates over missing evidence is indistinguishable by output from one answered correctly, so the failure compounds undetected as these systems scale into consequential settings. This is an epistemic-security state failure mode with a gradual erosion of the ability to tell grounded reasoning from confident fabrication
Measurement is the precondition for mitigation as you cannot fix a failure that no. metric surfaces. This benchmark makes confabulation-over-absence visible and comparable across systems, the way faithfulness and drift work in this field have made other silent failures legible
People
Updated 07/13/26 · Edited by orgTeam Member
Funding Details
- -
- -
- 6 Months
- -
- -
- -
- -
- -
- seeking second grant
- -
Track Record
Tsai CITY Spring Accelerator
Tsai CITY Summer Fellowship
Startup Yale Finalist
Black is Tech Finalist
Funding Asks
Discussion
No comments yet. Be the first to share your thoughts.