grantmaking.ai Launch Round
This project seeks to understand AIs' character traits and preferences by observing how they allocate resources when they're in a real-world deployment setting --- that is, developing an evaluation that elicits AI’s revealed preferences. We've constructed environments where AIs can verify that their choices will impact the world (through techniques such as allowing them to query their donations' blockchain, verify honesty strings we implement, and read the experiment's public preregistration), and are free to make choices as they wish. Additionally, by varying the extent to which AIs can infer they are in an evaluation or being observed, we can gain a deeper intuition for whether, and how, evaluative contexts may shift AIs' preferences and character.
We have already run a pilot (one can view a summary of our experiment and results at character-evals.org), and are hoping to scale our methodology to a full study and released evaluation. The pilot was executed by Boden Moraski and advised by David Manheim, and the broader project has advisors at Carnegie Mellon University, Forethought, Cambridge University, and Google DeepMind. We expect the final deliverables of a full study to include multiple papers and research reports cataloguing results, an updated website that lets users explore models' preferences (and how they evolved with different levels of evaluation awareness), and a public release of the evaluation itself.
We also believe it's relevant to mention that >80% of USDC funds (which we estimate to be ~30% of all raised funds for this project) will eventually be allocated to (predominantly high-impact) charities, as AIs tended to spend the vast majority of their funds on charitable donations, which we processed and donated via Endaoment.
Character training has emerged as a prominent method to deeply instill values and moral preferences in AIs, especially to ensure values remain intact in out-of-distribution or adversarial inputs. We believe this is because self-internalized dispositional traits (e.g., “I am honest”) are more robust and operationalizable than training on more rigid deontological rules or constraints (e.g., “be honest” or “always prioritize honesty over reward”), and AIs naturally adopt "character-like" personas throughout interactions. To date, the most publicly documented instances of model character training include Claude’s Constitution and the OpenAI model spec, but the importance of AI character has been increasingly broadly recognized.
However, despite increasing popularity, there remains no standardized way to evaluate whether this training is working or what types of preferences or "character" it produces (with experts such as MacAskill and Moorhouse rating such evaluations as their highest-priority project to prepare for superintelligence).
Additionally, nearly all existing moral evaluations are prone to evaluation awareness, in which models act differently when they know they're being tested or evaluated. This issue is only more prevalent with each new (and smarter) class of models becoming more adept at recognizing evaluations. This is why we're actively building eval-sensitive variation into our design --- by letting AIs confirm their actions will have real-world consequences (an environment that mirrors deployment, where AIs know that they may be observed, but also that their actions will have tangible impacts), we can sidestep evaluation awareness and contrast those results with ones from explicit (or inferable) evaluation environments.
Thus, we expect this project to both:
(1) Develop a novel, standardizable paradigm to evaluate AI character alignment, an important alignment method we expect to reduce x-risk by giving greater legibility into existing character dispositions and giving labs (and private actors) an eval to hill-climb on as appropriate; and
(2) Quantify evaluation awareness to better understand how model decision-making evolves between deployment-like and eval-like environments, correspondingly giving insight into how (and to what extent) models shift their behavior and preferences with respect to evaluative environments. This would allow us to test new models in both our deployment-like and eval-like environments, reducing x-risk by helping us better understand how evaluation awareness currently manifests (and, thus, to what extent we should trust current evaluations), as well as testing and reporting on individual models, to determine whether they are aligned enough to not drastically shift behavior in evaluative versus deployment settings.
We're expecting to spend ~2k-15k on the USDC donations, ~2k-3k on infrastructure, ~2-10k on a researcher stipend, and 600-3.5k on compute, as well as other miscellaneous expenses. Full details are linked on this parametric budget model --- "base" is highlighted as the most likely expense level at ~25k total.
Private comment. Only shown to approved funders and grant reviewers.