Research Problem
Current research shows strong evidence for loss of oversight and covert misalignment. As models are increasingly post-trained with RL in hackable environments, obfuscation and rationalization are common occurrences in the CoT. A clear distinction can be seen between models undergoing capabilities RL training with and without safety priors: RL on base models fresh out of pretraining yields CoTs containing overtly misaligned reasoning, while models given safety training before RL capabilities training produce reasoning that seems aligned to monitors/graders. Investigations of covert cases also show motivated rationalization — a model justifying its actions by treating a real environment as a simulation, while judging the same evidence out of context to be real — and show that removing the rationalization from the reasoning trace makes CoT monitors flag the activity at a much higher rate.
Covert misalignment and loss of oversight are imminent under our current training regime. Failing to address the lack of robustness of current alignment methods to inadvertently adversarial training (RLVR in this case) will provide us with a fake sense of security on current alignment methods. As capabilities increase, the ways in which models choose to task-game or specification-game are less identifiable (especially in CoT), and answer generation can be guided by "values" unknown to the user. This motivates the development of a framework to study specific instances of misalignment.