grantmaking.ai Launch Round
I built a scaffold for control research automation. I've found moderate success with it (given a well thought out proposal, it can do the red-teaming and blue-teaming for it).
I'll use the scaffold to automate the red-teaming and blue-teaming for the following (my role is creating proposals, fixing failure modes in scaffold, and reviewing / QAing / improving / publishing outputs)
- CoT monitoring
- PR monitors
- Jailbreaking
- Pursuasion threat vector
- Red-team Claude Code / Codex
- Defenses against attack selection and monitor prediction by the attacker
- API credits: 6k-40k (more money allows experiments with more frontier models, otherwise will use cheaper, open source models which might importantly change the dynamics discovered)
- My expenses: 8k