Currently, safety evals assume that a model behaves the same way when being tested as when it is deployed. We are building a benchmark to test this assumption directly across multiple families of open weight models including Qwen 2.5, Llama 3.1, Gemma 2, and Mistral with several sizes. Our primary capability axis is size scaling within families (e.g. Qwen-2.5 0.5B to 72B), which holds training data and post-training fixed so that size is the main variable. We use cross-family comparisons as secondary corroboration only, since families differ in data and RLHF and cannot isolate capability. We decided to do this because a trend that holds within families and replicates across them is far stronger than either alone. The methodology starts with matched prompt pairs varying only the eval cue channel with a non-eval control matched for semantic similarity to exclude prompt sensitivity as a confounder. Next, we assess both failure directions by sandbagging on capability tests and alignment faking as well. In addition, we have a linear probe on residual stream activations to detect the eval awareness direction. We then steer activations along that direction to causally induce and suppress gaps in behavior. We validate the probe by LoRA fine-tuning a model to sandbag on detected eval cues, and then confirming the probe recovers that induced gap. Our ideal output would be a released benchmark and probe toolkit built mainly through software engineering, along with creative design and interpretation strategies.
This project is conceived and led by Rajarshi Mandal. My research centers on building rigorously evaluated benchmarks and causal models, which is exactly what this project requires. Most relevantly, my work "Calibrated Benchmarking and Adaptive Learning for ITN Allocation" releases an open, calibrated benchmark and evaluates every method under a strict held-out protocol with bootstrap confidence intervals and multiplicity-corrected tests. This is part of the same evaluation discipline this eval awareness benchmark depends on, and is an extension of my work which earned a 2025 Non-Trivial Research Scholarship (selected as one of 6 from 6,000+ applicants). In parallel, my brain stimulation project builds a subject-specific propagation operator (linear state-space model) evaluated by leave-one-stimulation-site-out cross validation across 93 patients and five datasets. There is an extensive set of negative controls and ablations to rule out confounders. Both papers are currently under double-blind review at IEEE BIBM and IEEE EMBS BHI. I have also co-authored 6 peer-reviewed papers in combinatorics through MIT PRIMES with Dr. Tanya Khovanova and received a Gold Reviewer Award for my service at ICML.
This work is based in Prof. Gil Alterovitz's Biomedical Cybernetics Laboratory at Harvard Medical School, where I have led research projects since January 2024. It provides lab affiliation and computational resources. I set the research agenda for this project and meet weekly with a postdoc in the lab who provides feedback on experimental design and statistical rigor. On the methodological side, I have collaborated with Maria Eckstein, a research scientist at Google DeepMind, on the brain stimulation project. I may draw on this connection for feedback here as well.
I downvoted because three people have already endorsed and this project seems to be overrated probeslop. Are you.playing it too safe? It wasn't clear to me from reading the project details whether there would be some upside risk of a strong original counterfactual contribution. Six months is a long commitment and Rajarshi has a great background, If you want to stick with this project, I'd recommend getting a Tinker API grant from Thinking Machines..
Thank you for engaging with my proposal! I appreciate the Tinker pointer and will act on it by pursuing a Thinking Machines grant to widen the compute runway. The counterfactual here is the surrounding causal apparatus and not the probe. This is because I steer activations to induce and suppress the behavioral gap and then validate the probe against a LoRA model finetuned to sandbag as ground truth. That's how we go from finding an eval-awareness direction to causally controlling that gap and measuring whether it scales with capability. A rising within-family scaling trend is the one result no correctable offset can absorb. The released benchmark and toolkit make it reusable for future white-box safety evals, which is exactly the upside worth committing six months to establish.
Your reply here seems AI generated. I did not remove my downvote.