Run AgentHarm on 3–4 frontier/open-weight models, manually audit transcripts for metric gaming and spurious failures, compare to prior critiques, and publish a detailed evidence-based writeup.
Run AgentHarm on 3–4 frontier/open-weight models, manually audit transcripts for metric gaming and spurious failures, compare to prior critiques, and publish a detailed evidence-based writeup.
Project Details
Updated 07/12/26 · Edited by orgI will run the AgentHarm benchmark against 3-4 current frontier and open-weight models, reading the full transcripts to find where the benchmark's reported numbers diverge from what's actually happening.
Importantly to understand the cases where a model "passes" by gaming the metric instead of refusing harm, or "fails" for reasons unrelated to safety.
I'll compare my findings against known critiques (including recent work on the "owner-harm" blind spot in existing agent benchmarks, and findings about text-based refusal doesn't reliably transfer to tool-call behavior) to see whether I'm confirming good known.
The output will be a public writeup with detailed and step by step concrete transcript evidence.
This is also my first concrete project moving from my current background into AI safety evals work full-time as I transition into working in safe AI.
Theory of Impact
Updated 07/15/26 · By grantmaking.aiBenchmarks like AgentHarm are widely cited as evidence of model safety, but a benchmark that isn't regularly stress-tested becomes a proxy.
Careful, transcript-level critique is one of the cheapest ways to catch this before flawed benchmark results get used to justify deployment decisions or safety claims. My project aims to sharpen how much epistemic weight the field places on an existing, heavily-relied-upon one benchmark and to build my own judgment as an evaluator, which would contribute towards building my work on critiquing evals and going full time in Research in Benchmarks and Evals.
People
Updated 07/15/26 · By grantmaking.aiTeam Member
Funding Details
- Jul 26, 2026
- Aug 29, 2026
- 1.5 months
- -
- -
- -
- -
- -
- seeking first grant
- -
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.