I will run the AgentHarm benchmark against 3-4 current frontier and open-weight models, reading the full transcripts to find where the benchmark's reported numbers diverge from what's actually happening.
Importantly to understand the cases where a model "passes" by gaming the metric instead of refusing harm, or "fails" for reasons unrelated to safety.
I'll compare my findings against known critiques (including recent work on the "owner-harm" blind spot in existing agent benchmarks, and findings about text-based refusal doesn't reliably transfer to tool-call behavior) to see whether I'm confirming good known.
The output will be a public writeup with detailed and step by step concrete transcript evidence.
This is also my first concrete project moving from my current background into AI safety evals work full-time as I transition into working in safe AI.
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.