I have an automated control research scaffold. I'll use it to find the best control protocols to protect against various attacks and publish reports.
I have an automated control research scaffold. I'll use it to find the best control protocols to protect against various attacks and publish reports.
Project Details
Updated 07/24/26 · Edited by orgI built a scaffold for control research automation. I've found moderate success with it (given a well thought out proposal, it can do the red-teaming and blue-teaming for it).
I tested it here and created methodology for evaluating how good the automation is, making it a number go up game for safety.
I'll use the scaffold to automate the red-teaming and blue-teaming for the following (my role is creating proposals, fixing failure modes in scaffold, reviewing / QAing / improving / publishing outputs, and benchmarking how good the results are)
- CoT monitoring
- PR monitors
- Jailbreaking
- Persuasion threat vector
- Red-team Claude Code / Codex
- Defenses against attack selection and monitor prediction by the attacker
I will also benchmark how good the automation is compared to previous control research like:
Theory of Impact
Updated 07/30/26 · By grantmaking.aiSpeed up control research so we have better control measures that can catch scheming AI and reduce chance of high stakes failures.
People
Updated 07/30/26 · By grantmaking.aiTeam Member
Track Record
2nd author LinuxArena, a control arena used for risk evaluations at Anthropic.
Discussion
No comments yet. Be the first to share your thoughts.