I'm proposing a research agenda to map how a model's persona can be contaminated at different stages of the training pipeline. This matters more as models are increasingly trained on other models' outputs (e.g. synthetic data, model-based data filtering, post-training alignment processes like RLAIF / Constitutional AI) with less human oversight. We know from subliminal learning that model-to-model transfer is real, and often unexpected, and I’m worried about behavioral traits in particular that can transfer covertly and are hard to detect.
I'm prioritizing a first experiment on RLAIF and Constitutional AI (two widely-used alignment methods) in which I induce an undesired trait in the feedback model. This could be sycophancy, which reward models already tend to bias towards, or a malicious dark trait like spite. The key questions are whether that trait then emerges in the policy out-of-distribution – in unrelated contexts the feedback model never scored on – and whether the capability gap between the two models changes the degree of transfer.
I'll be the main person working on this, with mentors and peers from previous programs (e.g. ML4Good, SPAR, AI Safety Camp) to lean on for feedback and support. The concrete output is a preprint (extended to a conference paper), plus open source code.