Demonstration of emergent misalignment in markets of LLM agents
Demonstration of emergent misalignment in markets of LLM agents
Project Details
Updated 07/13/26 · Provided via application · VerifiedWe will use the money to run large-scale simulations that demonstrate emergent misalignment in multi-agent systems. This will build on our existing work that demonstrates LLM evolution towards incentives in simple markets (under review at PNAS).
There is burgeoning recognition that AI safety is a multi-agent problem, with several prominent position papers [1,2,3]. However, there is a gap in actual empirical demonstrations of multi-agent misalignment in realistic environments. Our understanding of emergent collective behaviour is very well developed in for example evolutionary theory [4], multi-agent systems[5], and psychology[6], with a general principle that behaviour is a function of both the individual and the environment. However, the traditional focus in AI safety is on individual alignment and not on the environment [1]. As such, the field would benefit from rigorous experiments that demonstrate that safety is an emergent property of both agents and environments.
The outputs will be:
- A paper analysing how and why misaligned behaviours emerge in market environments. We will target large multidisciplinary journals such as PNAS, Nature Machine Intelligence, or Science Advances.
- We will open-source the simulation code to allow others to also carry out experiments of these kinds, along with an accompanying paper explaining the methodology. I have experience publishing successful python packages (piecewise-regression.py [7]; over 100 stars on Github).
The research will be led by me (Charlie Pilgrim, Research Fellow, University of Leeds), with support from Richard Mann (Professor, University of Leeds), who is my PI in my current position.
References:
[1] Tomašev, N., Franklin, M., Jacobs, J., Krier, S., & Osindero, S. (2025). Distributional AGI safety. arXiv preprint arXiv:2512.16856.
[2] Hendrycks, D. (2023). Natural selection favors AIs over humans. arXiv preprint arXiv:2303.16200
[3] Hammond, L., Chan, A., Clifton, J., Hoelscher-Obermaier, J., Khan, A., McLean, E., ... & Rahwan, I. (2025). Multi-agent risks from advanced ai. arXiv preprint arXiv:2502.14143.
[4] Darwin, C. (1859). On the Origin of Species by Means of Natural Selection.
[5] Nisan, N., & Ronen, A. (1999). Algorithmic mechanism design. ACM STOC.
[6] Lewin, K. (2013). Principles of topological psychology.
[7] Pilgrim, C. (2021). piecewise-regression (aka segmented regression) in Python. Journal of Open Source Software, 6(68).
Theory of Impact
Updated 07/13/26 · By grantmaking.aiThere are several pathways by which this will reduce x-risk:
-
Paper on emergent misalignment in multi-agent systems. This evidence will influence researchers, policy makers, and AI organisations to take emergent multi-agent risks more seriously, and to include more rigorous safety protocols on individual model releases as well as regulation on the use of LLM models in markets.
-
Open source code and accompanying paper. The expertise we need to tackle multi-agent AI safety is found in researchers across evolutionary theory, economics, psychology, and mutli-agent systems. We don’t need to reinvent the theory or empirical approaches; we have been studying multi-agent systems for centuries. We need to give these researchers tools and show them how they can do good work in this new field.
-
My own career transition to AI safety. My background and skills give me a valuable perspective that is not usually found in AI safety (spanning mathematics, cognitive modelling, experimental psychology, complexity science, collective intelligence). This funding will help facilitate building my own credibility and publication record, which will then lead to further research opportunities.
People
Updated 07/13/26 · Edited by orgTeam Member
Discussion
No comments yet. Be the first to share your thoughts.