grantmaking.ai Launch Round
Project summary
A run-level reproducibility audit of published LLM safety-eval benchmarks. Phase 1 (PARROT) has been completed (5 models × K=5 runs × 1,302 items) replicating cleanly (Spearman 0.95–0.99): evidence that unreliability is metric-dependent, not blanket.
This grant funds phase 2: completing the SycEval audit, shipping the arXiv preprint (late August), and generalizing the audit method across 2–3 additional safety-eval benchmarks.
What are this project's goals? How will you achieve them?
Sycophancy benchmarks are a primary tool for detecting a core alignment failure: models telling us (researchers, evaluators, and users broadly) what we want to hear. If single-run confidence intervals conceal run-to-run variance, every safety case citing those benchmarks inherits that fragility. Goals: (1) SycEval audit complete, (2) arXiv preprint, (3) extend the audit harness to additional benchmarks with public replication reporting. The harness exists and replicates; the remaining work is execution.
Who is on your team? What's your track record on similar projects?
Independent researcher; I run this project solo, with academic supervision and capstone co-authorship from Dr. Abdulaziz Alharbi (GCU). Five years as an applied ML engineer at Deepgram building production speech and LLM training and evaluation infrastructure. MSCS capstone defended July 2026.
Currently in MATS Phase 2 review; Direct track record: the PARROT phase is complete and public, replicates cleanly (Spearman 0.95–0.99), and is already funded by a BlueDot Rapid Grant.
What are the most likely causes and outcomes if this project fails?
Most likely failure: the generalization phase finds benchmarks replicate cleanly, which is itself a publishable negative result and still de-risks the safety cases built on them. Operational risk is timeline slippage from solo-researcher capacity; mitigated by the funded runway being precisely what protects the research time.
How much money have you raised in the last 12 months, and from where?
$350 from BlueDot Impact (Rapid Grant, June 2026) for SycEval research compute on this project.
Safety cases increasingly rely on evaluation results. Labs use them to inform deployment decisions, RSP-style commitments cite them as evidence of acceptable risk, and third-party audits depend on them when assessing model behavior. But these decisions are only as reliable as the evaluations behind them. If a benchmark's conclusions change across repeated runs, then every safety case built on those results becomes less reliable. Despite this, relatively little systematic work has examined the reliability of safety evaluations themselves. Measurement infrastructure is rarely the most visible research direction, but it is foundational to every downstream decision that depends on benchmark results.
My own research is where I learned this matters. My capstone initially appeared to show that identity priming made open-weight models measurably more sycophantic. When I built a harness to replicate the result across runs, the effect disappeared: it had been a false positive caused by relying on single-run confidence intervals, not a real behavioral shift. Applying the same harness to PARROT produced the opposite outcome. The benchmark replicated cleanly across runs. The lesson wasn't that safety evaluations are broadly unreliable, but that reliability varies by metric, and we currently have no systematic way to distinguish robust evaluations from fragile ones before we build on them.
SycEval is where this failure mode should appear if it exists, which makes it the sharpest test of whether single-run instability is a real and general problem. This grant funds the work needed to answer that question rigorously.
Minimum (~$8k): ~$6k partial runway over 3–4 months to protect the research time, ~$2k compute to extend the reliability audit across 2–3 additional safety-eval benchmarks. The sycophancy preprint ships regardless; this floor funds the generalization beyond it.
Ideal (~$22k): ~$16k for one full protected quarter at a rate that offsets consulting income, ~$5k compute to run the audit across a broader benchmark set, ~$1k preprint production and misc. Sycophancy-specific compute is already covered by a BlueDot Rapid Grant, so this budget funds the generalization and runway, not the core sycophancy runs.