Project Details
Updated 07/22/26 · Edited by orgAdditional details have been added as of July 20, 2026; see "Additional experimental details" section below.
Since 2022, [1] a large body of existing work in AI safety has developed and applied the idea of language models as ‘persona simulators.’ [2] This view has been widely adopted for reasoning about models and interpreting safety-related empirical results, e.g. [3]. However, the safety community lacks both a formalization of the concept of ‘persona,’ and a specific predictive theory of how models represent and apply them. We think that these can be achieved, building on recent work describing language model behavior as Bayesian inference over latent concepts [4] [5] [6] and new interpretability techniques capable of discovering multi-dimensional and nonlinear features [7] [8] [9]. This project therefore aims to do three things:
- Provide a formal account of what a persona is in terms of observable behavior and what persona selection is in terms of in-context bayesian updates over models’ beliefs about such personas
- Operationalize that formal account in order to test it behaviorally in real LLMs, building on related work in toy models [10]
- Investigate the latent space geometry of personas and persona selection, building on related internals-based views of ‘belief state geometry’ [11] [12]
The goal of the project is to publish a paper by October either achieving these aims or else demonstrate evidence for why they are not achievable. Additional concrete details about experiments can be found in the 'additional information' section below.
I am currently working with Samy Mammeri and mentored by Thomas Jiralerspong. We have made some progress on (1) and (2). Funding would support me working full time on this project for 2.5 months (July-Sept).
References
[1] Janus, 2022, “Simulators” https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulators
[2] Marks et al. 2026, “The Persona Selection Model: Why AI Assistants might Behave like Humans” https://alignment.anthropic.com/2026/psm/
[3] Wang et al. 2025, “Persona Features Control Emergent Misalignment” https://arxiv.org/pdf/2506.19823
[4] Xie et al. 2022, “An Explanation of In-context Learning as Implicit Bayesian Inference” http://arxiv.org/abs/2111.02080
[5] Bigelow et al. 2026, “Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering” http://arxiv.org/abs/2511.00617
[6] Piotrowski et al. 2025, “Constrained belief updates explain geometric structures in transformer representations” http://arxiv.org/abs/2502.01954
[7] Wurgaft et al. 2026, “Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior” http://arxiv.org/abs/2605.05115
[8] Gurnee et al. 2026, “When Models Manipulate Manifolds: The Geometry of a Counting Task” http://arxiv.org/abs/2601.04480
[9] Bhalla et al. 2026, “Do Sparse Autoencoders Capture Concept Manifolds?” https://arxiv.org/abs/2604.28119
[10] Shai et al. 2025, “Transformers Represent Belief State Geometry in their Residual Stream” https://arxiv.org/pdf/2405.15943
[11] Bigelow et al. 2026, “Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space” https://arxiv.org/pdf/2605.12412
[12] Safari et al. “The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models’ Posteriors” https://arxiv.org/abs/2602.02315
Additional experimental details
UPDATE Jul 20, 2026
We have fleshed out additional experimental details. Some planned experiments include:
- Creating a cast of complex personas with shared and distinct traits, and confirming that methods based on [5] are able to predict models' exact probability distributions over discrete personas when in situations of ambiguity. This in itself is a significant step to supporting the persona selection model. These experiments will be done with both finetuning and in-context learning.
- Pretraining a very small LLM on this same corpus to determining the extent to which personas are 'privileged' in terms of the representations models learn, in the case of real text, à la [10]
Some more broad points about things we are thinking about:
- We think personas should be formalized as 'probability distributions over observed traits', with 'observed traits' being operationalized probably by
'average SAE feature activations on a judge model over many diverse rollouts; though SAEs have known pathologies, they are semantically diverse and quantitative, which is exactly what we need- We want to model not just the model's representations of personas, but its uncertainty about those personas/its probability distribution over personas; this is a part of the challenge
- Since the 'space of personas' is in principle continuous, we will start with thinking about discrete 'clusters of personas' with similar observed traits
- We can get ground truth about the models beliefs directly using logits in situations where the next token will resolve large amounts of uncertainty about a persona (e.g. a character saying a piece of information about themselves is true or false), giving us a specific empirical measure that internal geometry must predict.
- We're still figuring out the right experiments to run to develop a principled account of the internal belief-state geometry, but there's some related work we can draw on ([11]) to get started: using a dataset with large numbers of diverse characters interacting in diverse contexts in dialogue, we can extract activations on a base model and a post-trained judge model and use dimensionality reduction to get a manifold, then analyze the manifold to see if positional information is tied to persona information
Theory of Impact
Updated 07/22/26 · By grantmaking.aiThe persona-simulators view is widely applied in the safety community for reasoning about language model behavior. It makes consequential predictions about LLMs' sources of agency and goals (in particular, that only personas have agency and goals, and that there is no hidden 'shoggoth') and motivates both empirical safety training techniques (e.g. character training and inoculation prompting). Given how much work this paradigm is doing, refining it or refuting it would substantially improve the epistemic state of technical Al safety: if it is true, refining and formalizing it will enable more detailed accounts of whether and why current safety techniques work and thus the improvement of such techniques, and increase confidence in the specific technical x-risk research agenda that follows from the simulators view of language models. If it is false, understanding why it is false would help prevent risky or incomplete safety decisions from being made based on inaccurate intuitions. This project in mathematically formalizing and grounding the persona-simulators view in model internals would be a significant step towards validating (or falsifying) the view.
In addition to epistemic utility, a more accurate internals-based model of persona simulation would offer , since it would enable monitoring and control of model personas during post-training or deployment. (A demonstrative analogue would be [13], which demonstrated how reducing 'persona drift' can prevent certain harmful or strange behaviors and jailbreaks.)
People
Updated 07/22/26 · By grantmaking.aiTeam Member
Funding Details
- -
- -
- 2.5 months, plus potential extensions if successful
- -
- -
- -
- -
- -
- -
- -
Discussion
Hi @Logan Graves,
Excited to recommend the grant of $12,000!
Two quick questions before we distribute:
- Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
- Please confirm your commitment to post quarterly updates on how the project is going.
Hi Anton,
Delighted to receive the grant! Thank you for the recommendation.
(1) Funding situation has not changed — haven’t received any other funding and don’t need to update my funding ask.
(2) I will be happy to post quarterly updates on the project! I intend to have at least one paper out before the end of September, or else a blog post on why the hypotheses failed and related follow up work. Will that be good? I am also happy to provide informal/brief updates in some other forum if that’s useful.
Private comment. Only shown to approved funders and grant reviewers.