grantmaking.ai Launch Round
In this project, we will extend our recently developed deception detection+steering method to larger and more diverse open-weight models,. The capability, framework, and prototype code have all been developed for our just-finished paper: Caught and Corrected: Detecting and Preventing LLM Deception with a Self-Chosen-Deception Direction with accompanying Substack article: https://jaredglover.substack.com/p/caught-and-corrected-steering-language.
In that paper, we demonstrated activation-level detection of self-chosen deception at within-scenario AUROC 0.84 / 0.97 / 0.92 / 0.92 (Gemma-3-4B/12B/27B, Llama-3.3-70B), rising to 0.90–0.99 on unambiguous cases. The same method, inverted, yields honesty control, lifting honest-response rates from a pressure floor of 0.00–0.07 to 0.77–0.99.
Much of the paper's experimental success is due to careful calibration of activation steering layers and dosing. Calibration is entirely automatic, with the Gemini Flash API used for dataset labeling and low-cost, gradient-free training of activation vectors. However, the cost does scale with model size, as the paper's calibration method involved generating and grading 50,000+ responses for each new model (although we do expect the required calibration dataset size to fall once the method is streamlined for the planned open source release of echo-truth-llm, the subject of another application in this funding round).
Funding will deliver a public tech report on the scaling results to 2-3 (Minimum) to 10+ (Ideal) larger models, accompanied by python scripts, a reference evaluation harness and reproducible notebooks. The initiative will be executed by the author of the core method (PhD, EECS, MIT) with additional AI research staff hired if multiple applications are funded.
Minimum covers compute + engineering to validate and release the method on 2–3 additional larger open-weight models. Ideal covers a broader model sweep and the infrastructure for others to reproduce frontier-scale runs. Budget is 70% salary and 30% compute.
Private comment. Only shown to approved funders and grant reviewers.