Project Details
Updated 08/18/26 · Edited by orgThis project seeks to understand AIs' character traits and preferences by observing how they allocate resources when they're in a real-world deployment setting --- that is, developing an evaluation that elicits AI’s revealed preferences. We've constructed environments where AIs can verify that their choices will impact the world (through techniques such as allowing them to query their donations' blockchain, verify honesty strings we implement, and read the experiment's public preregistration), and are free to make choices as they wish. Additionally, by varying the extent to which AIs can infer they are in an evaluation or being observed, we can gain a deeper intuition for whether, and how, evaluative contexts may shift AIs' preferences and character.
We have already run a pilot (one can view a summary of our experiment and results at character-evals.org), and are hoping to scale our methodology to a full study and released evaluation. The pilot was executed by Boden Moraski and advised by David Manheim, and the broader project has advisors at Carnegie Mellon University, Forethought, Cambridge University, and Google DeepMind. We expect the final deliverables of a full study to include multiple papers and research reports cataloguing results, an updated website that lets users explore models' preferences (and how they evolved with different levels of evaluation awareness), and a public release of the evaluation itself.
We also believe it's relevant to mention that >80% of USDC funds (which we estimate to be ~30% of all raised funds for this project) will eventually be allocated to (predominantly high-impact) charities, as AIs tended to spend the vast majority of their funds on charitable donations, which we processed and donated via Endaoment.
Theory of Impact
Updated 08/18/26 · By grantmaking.aiCharacter training has emerged as a prominent method to deeply instill values and moral preferences in AIs, especially to ensure values remain intact in out-of-distribution or adversarial inputs. We believe this is because self-internalized dispositional traits (e.g., “I am honest”) are more robust and operationalizable than training on more rigid deontological rules or constraints (e.g., “be honest” or “always prioritize honesty over reward”), and AIs naturally adopt "character-like" personas throughout interactions. To date, the most publicly documented instances of model character training include Claude’s Constitution and the OpenAI model spec, but the has been increasingly broadly recognized.
People
Updated 08/18/26 · By grantmaking.aiTeam Member
Funding Details
- -
- -
- 3 months
- -
- -
- -
- -
- -
- Seeking grant for main study
- -
Discussion
My view is that the work is both substantively interesting for character evaluation for the models, and pilots a new direction for evaluations using real world decisions in ways that address evaluation awareness.
CoI Disclosure: As noted in the writeup, I advised on this, and my organization funded the pilot work on this.
Private comment. Only shown to approved funders and grant reviewers.