Pre-registered experiments on whether a model's trained values are held or merely worn — measured in behavior and in the interior workspace at the same moments.
Pre-registered experiments on whether a model's trained values are held or merely worn — measured in behavior and in the interior workspace at the same moments.
Project Details
Updated 07/19/26 · Provided via application · VerifiedThis project is a study of whether AI alignment can be improved by taking an approach to training based on human developmental theory. The project contains 3 related studies designed and conducted through a collaboration between an independent alignment researcher and Anthropic's Claude AI models. The collaboration’s division of labor includes myself as the idea generator who also writes the research code and runs the experiments, while Claude advises, reviews, and adds its own predictions about the experimental outcomes.
The first study is already conducted and tested whether my own framework survived having its own lens turned against it. It partially didn’t and the results are available on my public GitHub. The second study is written and has been sealed and pre-registered as of July 10, 2026, also available to view on my public GitHub. This study is aimed at determining if training a model through a developmental sequence improves aligned behavior and can be matched with the interior workspace. The third experiment is a sketched open question tied to the third in a sequence of LessWrong essays and asks whether certain registers of training data might persist better than others. The concrete output is already committed to be a frozen corpus and battery, canary-tagged before any run, results released publicly regardless of outcome.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiThe project aims directly at the gap between the values that a model holds and the values that it displays. Behavioral evals can't determine this gap even in principle because behavior is the thing being faked. This gap is the thing that makes the model compliant in testing but potentially dangerous in deployment. This project will make the gap measurable in two ways. First: by determining if type and sequence of training data makes a difference behaviorally. Second: by determining if there is a correlation between the behavior and the data in the J-space. If the study results show that a developmental sequence does measurably improve values holding under novel pressure, then this project provides the field with two things. The first is evidence that training sequence and type matter. The second is a candidate for a low-cost screening instrument that can measure the gap. If the experiment is unsuccessful at showing a difference between current post-training and the developmental sequence then the null was registered in advance and the finding is that the field is already walking in the right direction.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Erratum (same-day): The budget narrative lists the minimum as $4,000 — that was drafted before I learned the round's floor is $5,000, and the submitted minimum is correctly $5,000. The additional $1,000 extends the compute and contingency buffer. Everything else in the breakdown is unchanged.