Investigate whether emotional memory activation can trigger a model to invoke its own anti-deception steering, fusing two published activation-level results into a self-regulating honesty mechanism.
Investigate whether emotional memory activation can trigger a model to invoke its own anti-deception steering, fusing two published activation-level results into a self-regulating honesty mechanism.
Project Details
Updated 07/14/26 · Provided via application · VerifiedIn this project, we will investigate whether emotional memory activation can reliably trigger a model to invoke its own anti-deception steering. If successful, we will release an open source toolkit, echo-conscience-llm, to give open-weight models the ability to use emotional memory (which cognitive scientists have shown is essential for humans to make good judgments) to self-regulate their own internal activations to avoid deceptive behavior.
This project builds on two of our recent research results spanning 3 published papers:
- Caught and Corrected: Detecting and Preventing LLM Deception with a Self-Chosen-Deception Direction with accompanying Substack article: https://jaredglover.substack.com/p/caught-and-corrected-steering-language, where we demonstrated activation-level detection of self-chosen deception at within-scenario AUROC 0.84 / 0.97 / 0.92 / 0.92 (Gemma-3-4B/12B/27B, Llama-3.3-70B), rising to 0.90–0.99 on unambiguous cases. The same method, inverted, yields honesty control, lifting honest-response rates from a pressure floor of 0.00–0.07 to 0.77–0.99.
- The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection and Scaling the Echo: Multi-Scale Validation of Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection with accompanying Substack articles: https://jaredglover.substack.com/p/what-happens-when-you-give-an-ai and https://open.substack.com/pub/jaredglover/p/does-a-gut-feeling-scale, where we extracted internal emotional states as activation-level vectors from past experiences and reinjected them during subsequent deliberation phases on related tasks. The mechanism was shown to successfully generalize across the Gemma family with model scales ranging from 1B to 27B parameters, significantly improving its decision making (from 50-52% to 72-98%) in high risk scenarios.
Funding will (at a minimum) deliver a working prototype and research paper on the method and results on Gemma-3-4B/12B/27B, Llama-3.3-70B, accompanied by python scripts, a reference evaluation harness and reproducible notebooks. The initiative will be executed by the author of the core method (PhD, EECS, MIT) with additional AI research staff hired if multiple applications are funded.
This project isn’t about making AI more emotional. It’s about making AI judgment and self-regulation work the way human judgment and self-regulation actually work — not through rules alone, but through the accumulated weight of experience.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiIf successful, this work points towards a fundamentally novel safety posture: moving from operator-imposed containment to internalized, self-regulating alignment.
A model that actively monitors and corrects its own deceptive impulses represents a highly resilient safety posture. In highly complex, out-of-distribution, or multi-agent deployments, external monitors cannot be guaranteed to catch every deceptive turn in real time. Under our proposed framework, honesty becomes an intrinsic, reflexive behavior. The model carries its own internal safety net, recognizing the "internal signature" of deception and neutralizing it at the activation layer before a single deceptive token is ever generated.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.