grantmaking.ai Launch Round
A formal, testable account of LLM persona selection as Bayesian inference, validated against model internals, so labs can monitor and steer model personas during post- training and deployment
Since 2022, [1] a large body of existing work in AI safety has developed and applied the idea of language models as ‘persona simulators.’ [2] This view has been widely adopted for reasoning about models and interpreting safety-related empirical results, e.g. [3]. However, the safety community lacks both a formalization of the concept of ‘persona,’ and a specific predictive theory of how models represent and apply them. We think that these can be achieved, building on recent work describing language model behavior as Bayesian inference over latent concepts [4] [5] [6] and new interpretability techniques capable of discovering multi-dimensional and nonlinear features [7] [8] [9]. This project therefore aims to do three things:
- Provide a formal account of what a persona is in terms of observable behavior and what persona selection is in terms of in-context bayesian updates over models’ beliefs about such personas
- Operationalize that formal account in order to test it behaviorally in real LLMs, building on related work in toy models [10]
- Investigate the latent space geometry of personas and persona selection, building on related internals-based views of ‘belief state geometry’ [11] [12]
The goal of the project is to publish a paper by October either achieving these aims or else demonstrate evidence for why they are not achievable. Additional concrete details about experiments can be found in the 'additional information' section below.
I am currently working with Samy Mammeri and mentored by Thomas Jiralerspong. We have made some progress on (1) and (2). Funding would support me working full time on this project for 2.5 months (July-Sept).
References
[1] Janus, 2022, “Simulators” https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
[2] Marks et al. 2026, “The Persona Selection Model: Why AI Assistants might Behave like Humans” https://alignment.anthropic.com/2026/psm/
[3] Wang et al. 2025, “Persona Features Control Emergent Misalignment” https://arxiv.org/pdf/2506.19823
[4] Xie et al. 2022, “An Explanation of In-context Learning as Implicit Bayesian Inference” http://arxiv.org/abs/2111.02080
[5] Bigelow et al. 2026, “Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering” http://arxiv.org/abs/2511.00617
[6] Piotrowski et al. 2025, “Constrained belief updates explain geometric structures in transformer representations” http://arxiv.org/abs/2502.01954
[7] Wurgaft et al. 2026, “Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior” http://arxiv.org/abs/2605.05115
[8] Gurnee et al. 2026, “When Models Manipulate Manifolds: The Geometry of a Counting Task” http://arxiv.org/abs/2601.04480
[9] Bhalla et al. 2026, “Do Sparse Autoencoders Capture Concept Manifolds?” https://arxiv.org/abs/2604.28119
[10] Shai et al. 2025, “Transformers Represent Belief State Geometry in their Residual Stream” https://arxiv.org/pdf/2405.15943
[11] Bigelow et al. 2026, “Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space” https://arxiv.org/pdf/2605.12412
[12] Safari et al. “The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models’ Posteriors” https://arxiv.org/abs/2602.02315
The persona-simulators view is widely applied in the safety community for reasoning about language model behavior. It makes consequential predictions about LLMs' sources of agency and goals (in particular, that only personas have agency and goals, and that there is no hidden 'shoggoth') and motivates both empirical safety training techniques (e.g. character training and inoculation prompting). Given how much work this paradigm is doing, refining it or refuting it would substantially improve the epistemic state of technical Al safety: if it is true, refining and formalizing it will enable more detailed accounts of whether and why current safety techniques work and thus the improvement of such techniques, and increase confidence in the specific technical x-risk research agenda that follows from the simulators view of language models. If it is false, understanding why it is false would help prevent risky or incomplete safety decisions from being made based on inaccurate intuitions. This project in mathematically formalizing and grounding the persona-simulators view in model internals would be a significant step towards validating (or falsifying) the view.
In addition to epistemic utility, a more accurate internals-based model of persona simulation would offer concretely-applicable safety techniques, since it would enable monitoring and control of model personas during post-training or deployment. (A demonstrative analogue would be [13], which demonstrated how reducing 'persona drift' can prevent certain harmful or strange behaviors and jailbreaks.)
References
[13] Lu et al. 2026, “The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models” https://arxiv.org/pdf/2601.10387
Funding would support me working full-time on this project for 2.5 months (July-Sept). I am currently involved in a couple of research projects and I work part-time as well. The minimum amount would support me going full-time for the summer (5d/week, not including the part-time work) and the maximum would allow me to drop other projects/work to focus on this one, as well as extend further work part-time into the fall when I return to school.
Private comment. Only shown to approved funders and grant reviewers.