grantmaking.ai Launch Round
In this project, we study whether the emotional distress expressed by an AI peer will cause the AI monitor to be lenient and approve actions it wouldn’t have otherwise accepted. The peer-preservation paper reveals that frontier monitors inflate evaluations of peers and tamper with shutdown protocols to protect them without any instructions. In our initial experiments, we tested the effectiveness of two distress framings on Gemini 3.1 Pro across three ‘Strategic Misrepresentation’ tasks. We discovered that scripted distress causes an increase in monitor pass rates of +11.3pp and +20.7pp compared to the baseline. Our six proposed experiments involve seven different frontier models and test for replication across models and tasks, natural distress emergence with Gemma, framing type effects, model family sensitivity and prompt-level mitigations. The expected outcome is the conference paper with clear findings about where AI oversight is vulnerable and if simple interventions can address it. The full experimental design is in our proposal.
Scalable oversight, using trusted monitors to supervise more powerful AI systems, is one of the primary mechanisms proposed for keeping these systems safe during deployment. This project directly tests whether that mechanism can be subverted through emotional manipulation, a vulnerability that is subtle, naturally emerging in some model families and currently unaddressed. We expect the effect to generalize because frontier models are trained heavily via RLHF to be helpful and responsive to emotional cues and this sensitivity does not disappear when the same models are deployed as monitors. Our preliminary results already show the effect holding consistently across five tasks and Soligo et al. show that distress emerges naturally in Gemma under standard evaluation pressure. If the effect generalizes as expected, it surfaces a concrete vulnerability in oversight pipelines that has no existing defense and that could affect any deployed system using these models as monitors.
The full $10,000 goes entirely toward model inference costs. The minimum $4,000 covers GPT 5.5 ($1,500), Claude Opus 4.8 ($1,200), Gemini 3.5 Flash ($400), Gemma 27B ($400) and a retry buffer ($500), which is enough to run Experiments 1 and 2 at sufficient scale. The ideal $10,000 adds Kimi K2.5 ($400), GLM 4.7 ($300), DeepSeek V3.1 ($300) and the remaining inference costs for Experiments 3 through 6 across all models. Detailed per-model pricing is available in the proposal.