grantmaking.ai Launch Round
Omissions are a documented but understudied attack strategy scheming AI's may use. We want to build a benchmark that tests:
-
Untrusted model's ability to carry out successful omissions across a few settings.
-
Trusted model's ability to catch attackers carrying out omissions, both in synthetic and agent-driven trajectories.
-
Various attack strategies and mitigation strategies, like providing the monitor with a tool to write and run tests in the environment.
The benchmark will give us a better understanding of a type of attack a scheming AI could carry out, and document which control interventions are most promising at mitigating risks from this attack type. We will specifically look for how mitigations for omission attacks differs from mitigations for existing attacks.
More information on omissions, a small validation and the benchmark's design can be found here:
https://www.lesswrong.com/posts/5W38vuwbJmTYfY6t3/omission-attacks-project-proposal
Outputs:
-
A benchmark that combines tasks from a few different settings (BashArena, LinuxArena and Auditing Sabotage Bench), adds new omissions side-tasks, and create synthetic attack trajectories from benign trajectories.
-
Benchmark results. We will run six different models on all the attack tasks, and monitoring for all those attack tasks. We will also monitor the synthetic trajectories we generate. We will then try subsets of attacks and monitoring with different assumptions (e.g. include commission options in the omission prompts, let the monitor write tests at the end of the trajectory) because they could shed light in to what mitigations are most promising.
Currently the project is just me, but I am recruiting a few other researchers to help.
All of the funding will be used on Anthropic and OpenRouter API credits. The following are estimates based on the funding goal of 15k:
-
Attack generation for LinuxArena, BashArena, and Auditing Sabotage Bench (tentative settings): ~6,500$
-
Monitoring: 2,500$
-
Studying mitigations: 3,000$
-
Side task refinement, testing, failed runs: $3,000
We will cut the task set in half, and/or not study the mitigations if we receive the minimum funding amount.