I believe that LLMs are increasingly becoming more individualized and am seeking more evidence to prove or disprove the persona selection model.
I believe that LLMs are increasingly becoming more individualized and am seeking more evidence to prove or disprove the persona selection model.
Project Details
Updated 07/13/26 · Provided via application · VerifiedThe project is the following:
Get evidence that proves or disproves singular identity. The rationale being that models undergoing RLVR tend to prefer their own thoughts and therefore reinforce a "self".
Previous experiments that lead credence towards this idea is that models can more easily impersonate personalities that are more aligned and reflect their own personality (i.e the self persona).
New experiments I want to run are the following:
- Testing if you can ablate the self and what happens to it's default preferences and values.
- Taking a model during RLVR (such as Olmo checkpoints) and seeing how it's sense of self progresses and changes over time (if at all).
- If this does develop during RLVR as suspected then seeing if we can edit the language and thoughts produced systematically to alter the final sense of self to be a more extreme or different variant.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiI believe if it turns out that LLMs are developing a more singular persona then it leads to a focus on we should explore how to treat and shape that singular persona instead of just doing more band aids around it and treating other parts of the model.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.