Probe an open-weight model’s activations under biasing/cue conditions to test whether chain-of-thought explanations match internal reasoning, releasing paper, code, and datasets.
Probe an open-weight model’s activations under biasing/cue conditions to test whether chain-of-thought explanations match internal reasoning, releasing paper, code, and datasets.
Project Details
Updated 07/07/26 · By grantmaking.ai · VerifiedRecent developments in AI have led to Chain-of-Thought (CoT) reasoning becoming increasingly leaned on as an oversight signal. But does it faithfully reflect the model’s actual reasoning? Research from Turpin et. al and Chen et. al suggests that corrupting CoT or introducing cues can affect a model’s answer without being verbalized. I intend to combine activation probing with biasing an open-weight model to cross-check CoT and internal activations.
I’ll be working independently with the help of my mentor in Non-Trivial (selective research fellowship), expert feedback (connected with Martin Tutek already), and fellow peers. As an output, I expect a detailed paper, code repo, and released datasets so my results are reproducible. Any outcome, whether positive or negative, would be insightful in AI alignment: if we can trust CoT, we can more comfortably proceed with scaling AI models; if not, we can further investigate how to ensure CoT remains faithful by extending my results from varied task types.
I’m a student at Carnegie Mellon, so I know how important AI safety is. Moving forward in mech interp, I’ll be using knowledge gained from Leaf's Dilemmas and Dangers in AI course as well as Non-Trivial's fellowship program to guide my thinking.
Theory of Impact
Updated 07/07/26 · By grantmaking.aiLeading AI models can produce CoT reasoning as visible evidence of their reasoning. While this appears to be useful and reliable, what if this reasoning is generated after the model has already decided (i.e., post-hoc rationalization) rather than a genuine causal account of how it reached its answer? If that’s true, then a lot of oversight proposals that rely on reading AI reasoning would need to be revisited.
If we can ensure CoT reasoning is faithful, then we could see into the black box of AI reasoning. As long as reasoning stays faithful as AI models grow, alignment efforts would be aided by looking at CoT to ensure alignment. Current developments are already shifting towards CoT as a reliable source of oversight, but is this true? If my results are positive, then we can focus more on supporting CoT as an alignment tool; if results are negative, we could focus more efforts on other areas of AI alignment, like evals or scalable oversight.
People
Updated 07/07/26 · By grantmaking.aiTeam Member
Practically all of the Frontier Labs are betting on COT monitoring for safety and so faithfulness is an extremely important property to understand. However, 2K for a 3 month stipend seems to be on the lower end for this kind of work. I would encourage making that higher to enable a sustainable commitment.