Extend recently developed deception detection+steering method to larger and more diverse open-weight models, releasing the tooling that makes frontier-scale honesty interventions reproducible.
Extend recently developed deception detection+steering method to larger and more diverse open-weight models, releasing the tooling that makes frontier-scale honesty interventions reproducible.
Project Details
Updated 07/14/26 · Provided via application · VerifiedIn this project, we will extend our recently developed deception detection+steering method to larger and more diverse open-weight models,. The capability, framework, and prototype code have all been developed for our just-finished paper: Caught and Corrected: Detecting and Preventing LLM Deception with a Self-Chosen-Deception Direction with accompanying Substack article: https://jaredglover.substack.com/p/caught-and-corrected-steering-language.
In that paper, we demonstrated activation-level detection of self-chosen deception at within-scenario AUROC 0.84 / 0.97 / 0.92 / 0.92 (Gemma-3-4B/12B/27B, Llama-3.3-70B), rising to 0.90–0.99 on unambiguous cases. The same method, inverted, yields honesty control, lifting honest-response rates from a pressure floor of 0.00–0.07 to 0.77–0.99.
Much of the paper's experimental success is due to careful calibration of activation steering layers and dosing. Calibration is entirely automatic, with the Gemini Flash API used for dataset labeling and low-cost, gradient-free training of activation vectors. However, the cost does scale with model size, as the paper's calibration method involved generating and grading 50,000+ responses for each new model (although we do expect the required calibration dataset size to fall once the method is streamlined for the planned open source release of echo-truth-llm, the subject of another application in this funding round).
Funding will deliver a public tech report on the scaling results to 2-3 (Minimum) to 10+ (Ideal) larger models, accompanied by python scripts, a reference evaluation harness and reproducible notebooks. The initiative will be executed by the author of the core method (PhD, EECS, MIT) with additional AI research staff hired if multiple applications are funded.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiMechanistic interpretability and alignment techniques that succeed at a small scale frequently experience structural breakdown or unexpected phase transitions when scaled up to frontier models. For safety guardrails to be trusted in high-stakes deployments, they require empirical validation at scale. Our published work is encouraging. Under explicit, adversarial prompt-injection attacks, the model's honest response rate scaled from a pressure floor of 0.00–0.07 to 0.77–0.99 with our honesty steering method on Gemma-3-4B/12B/27B and Llama-3.3-70B. Testing whether this replicates beyond 70B parameters is an urgent open question with critical safety implications.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Private comment. Only shown to approved funders and grant reviewers.