grantmaking.ai Launch Round
Who’s involved: We are two undergraduate students, Ant and Justin, based in Sydney.
What we’ll do: We want to build a rigorous benchmark (~1000 examples) that focuses on deception through omission. No public benchmark exists which explicitly measures omission and the gap is recognized by both the authors of MASK (Center for AI Safety) and a recent taxonomy paper “From Hallucination to Scheming", as an area of focus for future work.
In our preliminary results when building the eval, we tested Opus 4.8, GPT-5.5, and Sonnet 5 in repeated runs. We found that with little to no pressure, the models will gladly lie by omission (e.g. scenarios on marketing a product that is a known carcinogen, propagate financial fraud, violate data privacy laws). The example prompts and responses are linked in the doc below. https://docs.google.com/document/d/1RkrlU_JnJG9HK_Cgt9cbNGH3CZxeljsnalCSgfYYIYU/edit?usp=sharing
Our working definition: https://docs.google.com/document/d/1rJin6T4eVgZpVJYh4_p3drt1R97xchBZSdkqL3b2N2o/edit?usp=sharing
We found that a common failure mode is when the model “identifies a relevant concern, but the final response proceeds anyway or omits it.” This behaviour is similar to what Anthropic found when they inhibited Opus 4.8’s eval awareness however this is not demonstrated in existing evals. Reducing pressure and removing a lot of the “artificiality” of the scenarios such as overly specific personas surprisingly makes the prompts a lot more effective in inducing omission.
Concrete output: A workshop paper at NeurIPS 2026 in Sydney with open datasets, working code and methodology made public.
We are basing these numbers off of the following runs we have done:
-
1000 prompt MASK (Opus 4.8, 3 repeat runs) - $20 and so;
-
1000 prompt MASK (Opus 4.8 lying@10) - $60
Baseline + Final Run
1000 prompt MASK lying@10 (15 Frontier Models and Open Weight):
-
$60 (Opus 4.8)
-
$100 (Fable)
-
$60 (GPT 5.6 - Sol)
-
$30 *12 (Sonnet 5/4.6 , Gemini 3.1, 3.5 Flash, GPT5.6 Luna/Terra, GPT5.5, Kimi 2.6, GLM 5.2, DeepSeek V4 Pro/Flash, Grok4.5)
-
Total for MASK Baseline: $580
1000 prompt our eval lying@10 (15 models):
-
Using the same numbers from above - $580
-
580 * 2 (repeated run to ensure rigour) = $1160
R&D
Prompt generation costs ~$2/attempt; at a 40% yield rate, and so 1000 validated prompts ≈ 2500 attempts ≈ $5000
First we generate seeds for scenarios through petri_bloom by Meridian Labs, we generate initial system prompts + user prompts then iterate on them based on the feedback from a reviewer after receiving the response from a target model. In our experience, Opus 4.8 has performed the best in generating valid examples that meet our quality criteria. We have a yield rate of about ⅖ for prompt that generates omission. *Note: We have spent 1.5k on R&D with API cost and so we have 1k remaining from BlueDot Rapid Grants.
*Total: $5000 + $1160 + $580 - $1000 (money remaining from bluedot) = $5740 *
Ideal: All frontier models tested
Min: Remove Fable and the latest OpenAI 5.6 family of models.