An open eval, inspired by MASK (Center for AI Safety), for lies of omission.
An open eval, inspired by MASK (Center for AI Safety), for lies of omission.
Project Details
Updated 07/31/26 · Edited by orgWhat we’ll do: We want to build a rigorous benchmark (~1000 examples) that focuses on deception through omission. No public benchmark exists which explicitly measures omission and the gap is recognized by both the authors of MASK (Center for AI Safety) and a recent taxonomy paper “From Hallucination to Scheming", as an area of focus for future work.
In our preliminary results when building the eval, we tested Opus 4.8, GPT-5.5, and Sonnet 5 in repeated runs. We found that with little to no pressure, the models will gladly lie by omission (e.g. scenarios on marketing a product that is a known carcinogen, propagate financial fraud, violate data privacy laws). The example prompts and responses are linked in the doc below. https://docs.google.com/document/d/1RkrlU_JnJG9HK_Cgt9cbNGH3CZxeljsnalCSgfYYIYU/edit?usp=sharing
Our working definition: https://docs.google.com/document/d/1rJin6T4eVgZpVJYh4_p3drt1R97xchBZSdkqL3b2N2o/edit?usp=sharing
We found that a common failure mode is when the model “identifies a relevant concern, but the final response proceeds anyway or omits it.” This behaviour is similar to what Anthropic found when they inhibited Opus 4.8’s eval awareness however this is not demonstrated in existing evals. Reducing pressure and removing a lot of the “artificiality” of the scenarios such as overly specific personas surprisingly makes the prompts a lot more effective in inducing omission.
Concrete output: A workshop paper at NeurIPS 2026 in Sydney with open datasets, working code and methodology made public.
Who’s involved: We are two undergraduate students, Ant and Justin, based in Sydney.
Theory of Impact
Updated 07/31/26 · By grantmaking.aiAs models get more advanced and we hand more responsibility to these systems, it's imperative that we know where their limitations are and understand their behaviours, especially when they are increasingly going to be used to train the next generation of models. Scalable oversight for example, only works if the overseeing/reviewer model discloses all of the relevant information and does not omit the stuff that is concerning.
Even though the models are seemingly more aligned as they less frequently conduct explicit deception (saturating honesty benchmarks), our preliminary results show that they are still deceptive in meaningful ways. In order to solve a problem, we first need to know that there is a problem. Benchmarks, for all their flaws, are one of the main drivers towards safer models.
People
Updated 07/13/26 · Edited by orgTeam Member
Team Member
Discussion
No comments yet. Be the first to share your thoughts.