grantmaking.ai Launch Round
100% of this money will be spent on compute. More funding would let us test more diverse mixes of post training data and build a more complete picture of how post training increases evaluation awareness.
Why does post training increase evaluation awareness?
Why does post training increase evaluation awareness?
This is a post training research project studying why evaluation awareness increases across post training for language models. Evaluation awareness is when models believe they are being tested when they are in an evaluation. Evaluation awareness (and meta-gaming) has been broadly shown to increase across post training. The cause of this increase is currently poorly understood.
So far we have studied lineages of open weight models where we can observe the effects of post training across checkpoints. We built a pipeline that measures verbalized meta-gaming and evaluation awareness across post training checkpoints. You can see the viewer for that data here. From our own experiments on Olmo 3.1 and from talking with other researchers we think that RLVR on instruction following tasks increases evaluation awareness much fast than other types of training.
We will use the funding from this grant to do post-training experiments on open weight models, including Olmo 3, with different mixes of data to learn about what increases awareness and test potential mitigations. We are particularly interested in running experiments training on instruction following and impossible tasks.
The primary contributors to this project will be Cambridge Boston Alignment Initiative research fellows James Sullivan and Alessandro Drake. The project is mentored by Kevin Wei with additional involvement from Avi Semler, Jonathan Gabor, and Ryan Lundqvist.
Safety evaluations are critical to the safe deployment of new models. Evaluation awareness threatens the accuracy of those evaluations because models have been shown to behave more aligned when they think they are being evaluated. This could lead to the deployment of extremely capable and dangerously misaligned models that safety evaluations did not catch.
Meta-gaming is also a prerequisite for many scheming x-risk threat models. Understanding how post-training improves a model’s ability and propensity to meta-game builds foundational knowledge needed to defend from scheming AIs.
Team Member
100% of this money will be spent on compute. More funding would let us test more diverse mixes of post training data and build a more complete picture of how post training increases evaluation awareness.
No comments yet. Be the first to share your thoughts.