A SCORE-style computational reproducibility audit of empirical AI safety research that estimates the field's base rate of reproducibility and generates a taxonomy of its failure modes.
A SCORE-style computational reproducibility audit of empirical AI safety research that estimates the field's base rate of reproducibility and generates a taxonomy of its failure modes.
Project Details
Updated 07/13/26 · Provided via application · VerifiedScience is an exercise in seeing further by standing on the shoulders of giants, and AI safety is now the field on whose shoulders the highest-stakes decisions stand. Lab training policies, government regulation, and hundreds of millions in philanthropic funding rest on empirical results that, today, no one systematically verifies. When we ran a reproducibility audit on the social and behavioral sciences (DARPA's SCORE program, reported in two 2026 Nature papers I coauthored), only 24% of papers even had the materials available to attempt a computational reproduction. No one has run this test on AI safety, a literature with added risk factors: much of it is published without peer review on proprietary models at a rapid pace.
In this project, I run the first systematic computational-reproducibility audit of empirical AI safety research. I have experience in building an automated agent-based system for performing these audits in the context of behavioral and social science research: a pipeline that, given a paper, finds its data and code, constructs the computational environment, re-executes the analysis, compares the outputs against what was reported, and compiles everything into a signed reproducibility package. Here, we adapt the work to the AI-safety corpus, which means handling everything from Alignment Forum posts without DOIs, to GPU fine-tuning workloads, to eval harnesses, and then applies the method to a random sample of empirical results, scaling with funding.
Theory of Impact
Updated 07/13/26 · By grantmaking.aiWhen the inputs and outputs of a field are not reproducible, the whole scientific ecosystem pays the price: in innovations delayed, resources wasted, investments misdirected, and trust eroded. For AI safety the price is denominated in existential risk. And regardless of which way an audit turns out, it is valuable, either for knowing what work can be relied upon, or for knowing which work needs firmer footing.
People
Updated 07/13/26 · Edited by orgTeam Member
Funding Details
- Sep 1, 2026
- Mar 1, 2027
- 6 months
- -
- -
- -
- -
- -
- -
- -
Track Record
We helped conceptualize and execute the DARPA SCORE program, which ran audits on hundreds of papers across the behavioral and social sciences, resulting in several papers on the topic, including:
- Miske, O., Abatayo, A. L., Daley, M., Dirzo, M., Fox, N., Haber, N., ... & Errington, T. M. (2026). Investigating the reproducibility of the social and behavioural sciences. Nature, 652(8108), 126-134.
- Aczel, B., Szaszi, B., Clelland, H. T., Kovacs, M., Holzmeister, F., Van Ravenzwaaij, D., ... & Geraldes, D. (2026). Investigating the analytical robustness of the social and behavioural sciences. Nature, 652(8108), 135-142.
We have also been designing automated systems to perform scalable audits of the computational reproducibility of research.
Private comment. Only shown to approved funders and grant reviewers.