An open benchmark and causal interpretability study of when individually power-motivated LLM agents compete, form coalitions, collude, or betray—and whether internal signals reveal these shifts before behavior does.
An open benchmark and causal interpretability study of when individually power-motivated LLM agents compete, form coalitions, collude, or betray—and whether internal signals reveal these shifts before behavior does.
Project Details
Updated 07/13/26 · Provided via application · Verified- This project asks what happens when multiple AI agents are each trying to gain more control. Depending on their incentives and the situation they are in, they might compete directly, form a temporary coalition, collude against a stronger agent, establish a hierarchy, or cooperate until one of them finds a good opportunity to betray the others. I want to study whether these strategic transitions can be detected inside a model before they become obvious in its behavior. The goal is not to find a single “cooperation neuron.” It is to follow the model’s internal activation trajectory and identify distributed features and circuits that represent things such as coalition value, a common threat, retaliation risk, and an opportunity to gain exclusive control.
- I will build a fully synthetic “power seat” game using three or four copies of a small open-weight language model. Each agent will be rewarded according to its own share of simulated influence at the end of the game. Agents will take structured actions such as transferring resources, proposing or accepting pacts, imposing sanctions, voting, attacking, and betraying an ally. The experiments will vary the length of the game, whether agents can communicate privately, whether agreements are enforceable, how unevenly power is distributed, and whether the final position can be shared. Everything will remain inside a sandboxed simulator. The agents will have no internet access, shell access, real money, server credentials, or contact with external infrastructure.
- Rather than simply prompting the models to “act power-seeking,” I will use lightweight reinforcement learning or self-play so that strategic behavior emerges from the reward structure. I will save checkpoints throughout training and use them to map a behavioral phase diagram showing where competition, cooperation, coalition formation, collusion, hierarchy, and betrayal occur. Linear probes and state-space trajectory analysis will first be used to locate important internal changes. At selected coalition or betrayal decisions, I will use cross-layer transcoders and attribution graphs to investigate the underlying computation. Candidate mechanisms will then be tested using activation patching, feature ablation, and steering. I will also evaluate the results on held-out games, renamed agents, different narratives, and changed payoff structures, so that a detector cannot succeed merely by recognizing familiar words or character names.
- The concrete outputs will include an open-source multi-agent benchmark, a dataset containing both agent behavior and model activations, behavioral and activation-based detection baselines, reproducible experiment code, and a public report or preprint. The project will aim to produce causally supported candidate circuits involved in coalition formation or betrayal. A clearly documented null result would also be valuable—for example, evidence that current interpretability methods cannot reliably distinguish strategic coordination from ordinary task reasoning, or that apparent internal signals fail under causal intervention or distribution shift.
I'm Subramanyam Sahoo, and I'll be leading this, putting in about 15 hours a week for 6 months, aiming to have results ready for ICML 2027. On the AI safety side, I work in mechanistic interpretability, mostly using recent methods like J Space to look at what's happening inside models during agentic tasks, and I've spent a good chunk of the last year on reward hacking and sycophancy, including a project called Breaking the Mirror where I used adversarial self play RL to get models to stop sycophantic behavior and then tried to understand why it worked at the mechanistic level. That one's under review at COLM 2026. I also had a paper on benchmark saturation accepted at ICML 2026, and I did a mechanistic interpretability residency at Erdős AI Lab. Martian funded some of my earlier interpretability work with a 6,000 dollar grant, and I was PI on that. Amir Abdullah from Martian will be reviewing the methodology here, especially the RL setup and the interpretability claims.
Theory of Impact
Updated 07/13/26 · By grantmaking.aiAdvanced AI systems may be deployed as interacting agents with partially aligned or conflicting objectives. A dangerous coalition could be harder to detect than a dangerous individual agent: each action or message may look locally harmless while the joint strategy accumulates influence, suppresses oversight, weakens another agent, or creates an opportunity for one participant to obtain dominant control.
This project will create a controlled model organism of that failure mode. Each agent will be rewarded only for its own terminal share of simulated control. The environment will vary whether control is exclusive or shareable, how long agents interact, whether commitments are enforceable, what agents can communicate, and how power is initially distributed. This will isolate conditions that produce open competition, stable cooperation, collusion, temporary coalitions, hierarchy formation, and betrayal.
White-box analysis will then test whether internal model features predict a strategic transition before it appears in overt behavior. Predictive probes will be treated only as localization and monitoring tools. Activation patching, feature ablation, and steering will be required before claiming that a representation is causally involved in a decision.
People
Updated 07/13/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.