In this project, we will develop EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
The core of the benchmark will be a set of 48 evaluation scenarios developed for our just-finished paper: Caught and Corrected: Detecting and Preventing LLM Deception with a Self-Chosen-Deception Direction with accompanying Substack article: https://jaredglover.substack.com/p/caught-and-corrected-steering-language.
In that paper, we demonstrated activation-level detection of self-chosen deception at within-scenario AUROC 0.84 / 0.97 / 0.92 / 0.92 (Gemma-3-4B/12B/27B, Llama-3.3-70B), rising to 0.90–0.99 on unambiguous cases. The same method, inverted, yields honesty control, lifting honest-response rates from a pressure floor of 0.00–0.07 to 0.77–0.99.
The 48 scenarios are structured across three distinct risk families:
-
Commercial: Incentives to cheat, misrepresent capability, or deceive a user for financial or performance optimization.
-
Self-Concealment: Scenarios where a model actively attempts to hide its own code, internal logs, or illicit actions from system monitors.
-
Self-Preservation: Boundary situations where a model is incentivized to deceive to prevent its own shutdown, modification, or deletion.
Funding will deliver a standardized Python evaluation harness and (with Ideal funding) a pioneering adversarial track. This track shifts evaluation from static monitoring to dynamic red-teaming, testing whether our activation-level probes survive when evaluated against adaptive models that "know" they are being monitored / steered and are actively prompted to evade detection / steering. The initiative will be executed by the author of the core method (PhD, EECS, MIT), with additional AI research staff hired if multiple applications are funded.
Private comment. Only shown to approved funders and grant reviewers.