grantmaking.ai Launch Round
CaML researches how to use midtraining (/Synthetic Document Finetuning) to make AI systems morally open-minded and compassionate toward all sentient beings. In this project, we’ll be researching how our midtraining data can be scaled so that the desired effects are not degraded by later fine-tuning.
Our previous work includes:
-
Releasing benchmarks on Inspect (including TAC, MORU and ANIMA) evaluating moral reasoning under uncertainty and compassion in frontier models
-
Conducting research showing how compassion can be effectively instilled during midtraining, but that fine-tuning will sometimes erode desired values
-
Finding that broad compassion for non-humans generalises to humans
Midtraining may produce deeper values (e.g. here) without the risks of RL, suggesting it is likely to generalize better than finetuning to transformative AI and is already used in frontier model pipelines. But we lack research into how different aspects of midtraining are affected by scaling, or how self-fulfilling alignment can be used. Some specific questions we’re looking to answer in this project include:
-
How does scaling midtraining data affect performance on our TAC benchmark for given post-training?
-
What mechanisms degrade compassion in finetuning?
-
How does scaling the fine-tuning (SFT, RLAIF, RLVR) and (compassion) midtraining affect persistence of the chosen value?
-
How do these effects scale with larger open-weights models?
-
Investigate how this self-fulfilling alignment interacts with self-other overlap techniques (with John Lund and Sentient Futures)
We’ll publish the results of these findings through arXiv. If the work demonstrates promising mechanisms for scaling compassionate midtraining, we’ll seek to present it directly to labs, and release a scaled dataset for the alignment community.
$50k
We’d investigate both self-fulfilling alignment and the effects of fine-tuning on scaled midtraining datasets. We could enrich our methodology, create more data, and invest more resources in communicating our findings to labs and the alignment community.
-
55% contractor pay (~750 hours spent on research and supervision of volunteers over ~20 weeks).
-
40% compute (as above, spending primarily on testing larger models (e.g. ~400b open-weights models with more data to provide stronger evidence of scaling to labs)
-
5% research communication (event travel, writing blog posts, lab outreach).
CaML has pioneered a lot of the novel research around technical AI safety for nonhuman welfare, and I'm looking forward to seeing what new outputs they produce.