grantmaking.ai Launch Round
In this project, we will develop echo-thruth-llm, an open source toolkit to detect and correct deception in open-weight models. The capability, framework, and prototype code have all been developed for our just-finished paper: Caught and Corrected: Detecting and Preventing LLM Deception with a Self-Chosen-Deception Direction with accompanying Substack article: https://jaredglover.substack.com/p/caught-and-corrected-steering-language.
In that paper, we demonstrated activation-level detection of self-chosen deception at within-scenario AUROC 0.84 / 0.97 / 0.92 / 0.92 (Gemma-3-4B/12B/27B, Llama-3.3-70B), rising to 0.90–0.99 on unambiguous cases. The same method, inverted, yields honesty control, lifting honest-response rates from a pressure floor of 0.00–0.07 to 0.77–0.99.
Much of the paper's experimental success is due to careful calibration of activation steering layers and dosing. Calibration is entirely automatic, with the Gemini Flash API used for dataset labeling and low-cost, gradient-free training of activation vectors. Thus, once the first version of echo-thruth-llm is released, it will be straightforward to add support for a wide range of other open source models beyond the initial Gemma-3-4B/12B/27B and Llama-3.3-70B.
Funding will deliver a concrete, pip-installable library under a permissive MIT/Apache license, accompanied by a reference evaluation harness and reproducible notebooks. The initiative will be executed by the author of the core method (PhD, EECS, MIT), with the ideal budget adding support for more model families, HuggingFace transformers integration, and LangChain wrappers.
Minimum budget covers ~3 months part-time to package, document, and release the detection+steering pipeline for Gemma-3 (4B/12B/27B) and Llama-3.3-70B, with reproducible notebooks and the eval harness. Ideal budget adds support for more model families, HuggingFace transformers integration, and LangChain wrappers.