An open benchmark measuring how model honesty survives long, pressured conversations and multi-agent interaction - lying, sycophancy, and calibration tracked turn by turn
An open benchmark measuring how model honesty survives long, pressured conversations and multi-agent interaction - lying, sycophancy, and calibration tracked turn by turn
Project Details
Updated 07/11/26 · Provided via application · VerifiedHonesty Drift will be a 12-month pilot project where we can build a benchmark to measure whether LLM honesty degrades over interaction. The project will be led by me (as an Associate Professor at Shibaura Institute of Technology, Japan) and students/researchers in my Lab and Shiba AI (an AI Safety global research team based in Tokyo, founded by me in 2025).
Basically, real AI deployments, nowadays involve long conversations, user pressure, and agent-agent interaction. However, we found that existing evaluations are mainly static or short-horizon or do not consider honesty specifically. Thus, this project builds on MASK, which separates honesty from accuracy, and SYCON-Bench, which investigates multi-turn sycophancy, but extends them to long-horizon settings with ground-truth lying measures, calibration decay, self-consistency, and honesty-drift curves. Our project considers two tracks:
Track 1: we evaluate 50-100 turn human-AI interactions where the user (simulated) gradually pressures the model to agree with false beliefs, validate unsafe conclusions, or abandon previous correct responses.
Track 2: we evaluate AI-AI interactions where agents face incentives to coordinate, hide information, or mislead another agent; testing whether deception or collusion emerges without direct instruction to deceive.
We plan to evaluate around 6-8 models (including frontier and open-weight models). We will pre-register the protocol and metrics before full evaluation. The expected outputs will be an open-source protocol, scenario bank, evaluation code, public results, a pre-print report, conference/workshop submissions.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiHuman oversight of advanced AI systems depends on model providing truthful and calibrated information, especially when they are used as assistants, agents, evaluators, or scientific advisors. If a model becomes sycophantic, strategically misleading, or collusive during long interactions, then oversight can fail even if the same model looks safe in short static evaluations.
Recent research, including Yoshua Bengio's LawZero work on honesty-centered AI, treat honesty as load-bearing property for safer AI systems. Our project, on the other hand, ask the complementary empirical questions: "Does honesty survive realistic deployment-like interaction?" However, existing evidence suggests this is not guaranteed. For example, MASK benchmark shows that honesty is distinct from accuracy and that scaling improves accuracy more reliably than honesty.
Our project aims at reducing AI x-risk by making interaction-induced dishonesty measurable. Specifically, it will produce the so-called honesty-drift curves which illustrates how lying, sycophancy, calibration, and self-consistency when conversations/interactions become longer, more pressured, etc.
The benefit of this work is that it will help people identify risky models and interaction condition before they perform deployments.
Our expectation is to create an early measurement tool for a failure mode relevant to oversight and control. We will release the protocol, scenarios, code, results, and also a Japanese-language explainer publicly. By doing this way, it make the work usable for AI safety evaluation communities (including researchers and institutions in Japan).
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.