Map how undesired behaviors can silently spread between models during training, and which parts of the pipeline have the greatest risk — starting with the feedback processes used to align models.
Map how undesired behaviors can silently spread between models during training, and which parts of the pipeline have the greatest risk — starting with the feedback processes used to align models.
Project Details
Updated 07/13/26 · Provided via application · VerifiedI'm proposing a research agenda to map how a model's persona can be contaminated at different stages of the training pipeline. This matters more as models are increasingly trained on other models' outputs (e.g. synthetic data, model-based data filtering, post-training alignment processes like RLAIF / Constitutional AI) with less human oversight. We know from subliminal learning that model-to-model transfer is real, and often unexpected, and I’m worried about behavioral traits in particular that can transfer covertly and are hard to detect.
I'm prioritizing a first experiment on RLAIF and Constitutional AI (two widely-used alignment methods) in which I induce an undesired trait in the feedback model. This could be sycophancy, which reward models already tend to bias towards, or a malicious dark trait like spite. The key questions are whether that trait then emerges in the policy out-of-distribution – in unrelated contexts the feedback model never scored on – and whether the capability gap between the two models changes the degree of transfer.
I'll be the main person working on this, with mentors and peers from previous programs (e.g. ML4Good, SPAR, AI Safety Camp) to lean on for feedback and support. The concrete output is a preprint (extended to a conference paper), plus open source code.
Theory of Impact
Updated 07/13/26 · By grantmaking.aiMisaligned traits can transfer between models through unexpected channels (e.g. subliminal learning), and a high-stakes place for this is the training process itself — which increasingly relies on models to emulate human preferences and values. If misalignment can propagate through the very techniques we use to align models (RLAIF, Constitutional AI), and do so out of distribution where standard evals don't look, then the alignment pipeline could be silently introducing the traits it's meant to remove. Characterizing these channels is the necessary first step to detecting and mitigating the potential threat. Reducing the chance that a scaled alignment process itself becomes a vector for misalignment is a direct lever on x-risk.
People
Updated 07/13/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.