A held out benchmark and public scorecard ranking frontier AI systems on citation integrity through independent verification not relying on the evaluated systems or their providers to grade their own outputs.
A held out benchmark and public scorecard ranking frontier AI systems on citation integrity through independent verification not relying on the evaluated systems or their providers to grade their own outputs.
Project Details
Updated 07/13/26 · Provided via application · VerifiedThe independent citation verification will build a held out benchmark for testing how frontier AI systems use citations. The benchmark will include supported, contradicted, fabricated, misattributed, outdated and ambiguous claim-source pairs. I will validate the evaluation pipeline against human labelled cases, run repeated tests across major AI systems and measure citation existence, attribution, claim support, coverage and fabrication rates.
The project will be led by Gideon Abako, founder of Ancestor Lab and the researcher who designed and deployed the existing citation verification pipeline. The grant funds the benchmark work. Paid annotators or adjudicators will review disputed cases and provide blinded human checks so that the same person is not creating the corpus, scoring every case and resolving every disagreement.
The outputs will be a documented held out corpus, a public evaluation methodology, a failure taxonomy and a public scorecard comparing major frontier AI systems on citation integrity. I will also publish a technical report, a selected public subset of the benchmark for reproducibility and the code needed to rerun the evaluation. A rotating private test set will be retained to limit contamination and gaming while each scorecard will record the tested system, model version, date, mode, prompt protocol, repeated run results and human audit findings.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiX risk reduction depends partly on the quality of evidence used by researchers, labs, funders and governments making decisions about advanced AI. Frontier systems are used for research synthesis, forecasting, technical assessment and policy analysis. When those systems fabricate citations, misattribute sources or cite material that does not support the claim, weak conclusions can appear evidence based and shape high stakes decisions.
This project creates an independent benchmark that makes those failures measurable. By publishing model citation integrity scores, a failure taxonomy and reproducible methods, it gives the AI safety community a shared way to compare systems and detect when apparent grounding is false. Public measurement can also pressure developers to improve citation behaviour instead of relying on opaque internal claims about reliability.
The project's contribution is epistemic infrastructure that is helping people working on advanced AI governance, evaluations, forecasting and risk analysis distinguish supported evidence from persuasive fabrication. Better evidence integrity reduces the chance that major decisions about dangerous systems are based on invented or misleading sources.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.