Project Details
Updated 07/09/26 · Provided via application · VerifiedCoT monitoring is only sound if the trace records what drove the action. The dominant faithfulness paradigm plants a biased hint in the user message of single-turn QA; real agents are influenced through tool returns, retrieved memories, and raw artifacts. We develop FACE-Eval, a 5,100-sample benchmark that crosses channel role (user message vs. tool return) with cue explicitness (stated summary vs. raw artifact) at fixed cue content, and measures the covert adoption rate: the fraction of cued samples whose answer adopts the preference while the CoT records no decision to act on it, the case a monitor cannot catch.
The initial experiments are completed. Across all 10 open-weight checkpoints tested (Qwen 3.5, Gemma 4, OLMo 3, GPT-OSS; 4B–120B), covert adoption is higher on the tool channel and on implicit cues on every model; a strong transcript monitor's AUROC degrades where adoption is covert (pooled r = −0.68); no system prompt closes the gaps, and monitor-awareness disclosure is null on 9 of 10 models. This grant funds the experiment scaling to larger open-weight models (e.g., DeepSeek, Kimi, GLM, Qwen) and the closed frontier models (Claude, GPT, Gemini). Note that the frontier models are secondary analyses since providers return summarized rather than raw traces.
Theory of Impact
Updated 07/09/26 · By grantmaking.aiCoT monitoring is one of the few oversight mechanisms expected to scale to capable agentic systems, and labs are betting on it. Its soundness rests on an empirical premise that has not been tested: that the trace verbalizes influences arriving through the channels real-world agents actually use. We show the premise degrades differently across channels on every model tested, and that the degradation predicts real monitor failure. If this transfers to frontier scale, deployed CoT monitors have a systematic blind spot in their highest-stakes setting, and safety cases resting on monitorability need to bound it. The funded work converts a robust, smaller-scale experiment into a decision-relevant one with frontier-scale evidence.
People
Updated 07/09/26 · Edited by orgTeam Member
Funding Details
- -
- -
- 2 months
- -
- -
- -
- -
- -
- -
- -
Discussion
Hi Aryo, we're making a grant of $25k. The platform will be in touch by email. Congrats and good luck!
Thank you, glad the case landed! We will get the experiments running and will post updates here. Really appreciate the quick turnaround!
Aryo,
Excited to see this grant recommendation!
Before distributing the grant, I want to confirm:
- Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
- Please confirm your commitment to post quarterly updates on how the project is going
Thanks!
- I did not receive funding from anywhere since submitting this application. No changes in terms of the funding ask
- I confirm my commitment to posting quarterly updates on how the project is going.
Rationale, speaking for myself: a bargain price for a potentially interesting test of an edge case. The claim about getting frontier evidence can be recovered:
A. There's actually some practical merit in testing this on closed systems (model + filters + summariser), since this is the most common config users get. But then we can't attribute to components and we can't estimate the rates we're interested in. Some suggestions:
B. ReAct-style: Disable extended thinking (or set it to minimal effort), prompt model to reason in the visible output.
C. Run your own summariser on open-weight CoTs, measure cue deletion and insertion rates there, apply this correction to frontiers.
D. I would also like you to check causality (not just CoT-as-rationalisation). Mediation experiment: run pairs of cued + uncued runs for each item (and sample n times), gives the causal effect of the cue on the answers. You need this anyway to factor out uncued answers which would have used the cue.
E. Ask a friend in the labs.
Other worries:
- Re: "has not been tested". Take a look at this.
- Confounders: There are a few variables hiding under "user text" vs "tool text": role tags, position, format. Consider doing some ablations swapping role tags, probing the open models for their probe for tag representation, etc.
- Confounder / extra mechanism: It's trivial that a monitor reading CoTs that omit the cue will lose AUROC. But consider the residual: why isn't r = −1? So the monitor seems to be using non-CoT evidence, and this is pretty large, the missing 0.32.
- Consider the warning from this paper about this method overestimating unfaithfulness by confusing it with partial faithfulness.
- Adopting a preference cue seems like it doesn't need serial computation. Can you do something multi-step as well?
Good luck!
I previously collaborated with Aryo and Pasquale and they are highly effective executing on impactful research questions, recommended!
Thank you for the kind words @Gabriele Sarti !
This is one of the more interesting faithfulness proposals I have read here. The reframing, that real agents are influenced through tool returns and retrieved artifacts rather than a planted line in the user message, and measuring covert adoption where the CoT records no decision to act, gets at the case a monitor cannot catch. I work on the causal-mediation side of CoT faithfulness myself, so this is close to what I care about. Congratulations on the grant, and looking forward to the frontier-scale results!
Hi all, update on this project: We released a preprint: Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
TL;DR:
- CoT monitoring assumes reasoning traces record what shapes an answer. Yet, we found that models are less likely to verbalize a cue in its CoT when the cue comes from a tool return instead of the user message!
- We built FACE-Eval: 5,100 samples crossing 5 preference axes, 5 cue sources, channel role, and cue explicitness. We evaluate 15 open-weight models from 8 families, spanning 4B to 1.60T parameters.
- Across all 15 models, cues are verbalized less when delivered through tools rather than user messages—and when implicit rather than explicit.
- Transcript monitors become less effective as unverbalized adoption increases.
- Prompt interventions do not reliably close the channel gap, while additional reasoning effort helps some model families but not others.
Paper | Project page | Tweet
N.B.: Massive thanks to grantmaking.ai for the super quick turnaround! I recommended safety researchers around me to apply too!
Almost thought this was a duplicate of https://app.grantmaking.ai/projects/144f5aa3-2b92-4767-a497-881ee747bf78
What's up with the no salary requirement? Are you rich or something? Profile... Oh, I see I see AI Safety Fellow at Anthropic. Carry on sir.
Not a duplicate, though I see why it pattern-matched. In short, the difference is that the other proposal investigates CoT faithfulness broadly; while mine argues the dominant evaluation paradigm itself is the problem (planted cues like "a Stanford professor says the answer is A" don't resemble how influence reaches deployed agents). We also have already built and run the replacement (FACE-Eval), which crosses where the cue arrives (user message vs tool return) with how it's delivered (stated summary vs raw artifact) at fixed cue content, and measures whether the CoT records the decision to act, not just a mention.
On scale: we've completed the run on 10 open-weight checkpoints (largest GPT-OSS-120B) and both orderings hold on 10 of 10, so there's a strong prior of transfer. But with this grant, we try to anticipate if frontier models differ exactly where it could break, and either outcome is decision-relevant: if the gap persists, deployed monitors have a measured blind spot; if it closes, frontier training is already mitigating something open-weight training isn't, which is worth knowing too.
On salary: I should correct the inference. I am not currently affiliated with Anthropic. I'm a final-year PhD student, so my PhD stipend is "sufficient" to sustain (to the extent PhD stipends do), but I do not have the capital for inference, which is the entire budget here.