grantmaking.ai Launch Round
We are replicating, stress-testing, and extending main experiments in Teaching Claude Why, in which Anthropic introduces new alignment techniques used in current production Claude models. In the post, Anthropic uses synthetic document fine-tuning (SDF) and supervised fine-tuning (SFT) to align and attach the model to the "Claude" character, and finds that this technique successfully mitigates agentic misalignment propensities, with the effect persisting through reinforcement learning (RL) post-training.
Additional funding would allow for uncertainty measurements on our key findings, experiments testing persistence in long RL runs, and scaling laws for SDF corpus size. Our current funding is only sufficient for delivering a set of minimal viable products, and we think persistence testing is very important for frontier labs that are currently doing long RL runs.
Anthropic’s key claims are:
-
Training on documents about Claude's constitution and fictional stories about AI behaving admirably improves alignment despite being out-of-distribution (OOD) of the agentic misalignment evals.
-
Training on demonstrations of desired behavior is often insufficient; explaining why some actions are better than others matters, as does training on richer descriptions of Claude's overall character.
-
Data quality and diversity are crucial.
At a high level, we would like to test the following claims on an open-source model:
-
Is alignment tied to the AI assistant persona specifically? Or, is it more dependent on the quality and generalizable principles illustrated in the content of the data? That is, does the data need to be tied to the AI persona, and how much does the Personal Selection Model apply?
-
How important is the reasoning for instilling moral decision-making?
-
Do improvements from SDF persist through alignment RL and/or capabilities RL?
Our theory of change for this research is to provide an open-source replication to verify the paper’s claims, and for the broader alignment research community to inspect and build on. Inventing better alignment techniques is out of scope for this project, but we believe that our work could inspire them.
This project is being conducted as part of the Second Look Research fellowship at UChicago XLab, focused on replicating load-bearing results in AI safety. Research is primarily conducted by the fellows with oversight from the core team.
Who is involved:
-
Anastasia Wei (leading the project, will be first author)
-
Stewy Slocum (mentor, xAI)
-
Arav Dhoot, Finn Cairns, Jack Thompson, Brandon Qi (co-authors)
-
Yixiong Hao, Zephaniah Roe (research management)
-
Harshul Basava (operations/logistics)
Anthropic’s Alignment Science’s Teaching Claude Why blog describes the alignment-training techniques behind its current production models and proposes principled interventions for alignment mid-training. These are among the few alignment techniques shown to be practical and effective in production so far; whether they are robust, generalize OOD, and persist through RL has direct impacts on whether other frontier labs should adopt similar practices.
Anthropic's current alignment plan is implementing mid-training with constitutional documents (roughly "Teaching Claude Why"), and it’s plausible that they will then hand off alignment research to (hopefully) aligned human-level AIs. Since mid-training alignment is central to Anthropic’s strategy for aligning superintelligence, we must verify that these methods are robust. Otherwise, recursive self-improvement could easily amplify minor alignment failures into catastrophic outcomes.
Regardless of the outcome, our work will produce useful signals. If the mid-training step works well in an open-source model, it could encourage other frontier labs to adopt similar techniques. If it works but for unexpected reasons, it would help us and other labs to design more robust techniques. And if it fails completely, it could point to execution errors on our end or suggest these techniques don’t easily transfer to different model families. In this scenario, we will focus on disentangling why the technique fell short.
We plan to open source all of our code and models and their training checkpoints, enabling the broader AI safety community to pursue further research in:
-
Extension of mid-training alignment techniques
-
Data-centric interpretability research
-
Interp work on what mid-training instills in models
-
Associated claims about the personal selection model
Because some experiments are quite expensive, it is unlikely that other groups prioritize running multiple seeds and have statistical tests. We also have not seen concrete demonstrations from frontier labs that these methods scale in long RL runs, and we aim to verify the robustness of these mid-training techniques with the additional funding.
For 38-60k, we expect to accomplish the highest priority items in our proposal if no complications arise. 70k would provide a buffer for resolving unforeseen issues around data generation and leave no financial bottlenecks for the ideal version of our project. See budget breakdown linked here.
The “current budget” column denotes the funds we have already allocated to the project in USD from Second Look. The next two columns indicate our funding asks to improve the quality and impact of our work.