Project Details
Updated 07/14/26 · Edited by orgI am a PhD student at USC. I would like to develop benchmarks that can empirically distinguish between the different models of reasoning and behavior that LMs may be emulating. These benchmarks will target different categories of model behavior such as (a) how the consequences of an action are evaluated (e.g., causal decision theory vs evidential decision theory vs functional decision theory), (b) decision making under uncertainty (e.g., expected utility theory vs cumulative prospect theory), (c) strategic beliefs (e.g., quantal response equilibrium vs level k reasoning vs cognitive hierarchy models) and (d) outcome utility (e.g., Fehr–Schmidt vs Charness–Rabin).
Concretely, I would like support to cover the compute costs of the current and next stage of this project. Currently, I am implementing and evaluating an initial set of benchmarks aimed at examining how the behavior of language models in economic games changes when they learn that their opponents are copies of themselves. This work builds off prior work by Za et al. (https://openreview.net/forum?id=q949uxeGal) by examining how this knowledge affects: (a) higher order beliefs about group behavior in settings such as the p-Keynesian beauty contest and minimum effort game, (b) their ability to select between symmetrical equilibria in settings such as Hawk–Dove, the 11–20 game, and the Battle of the Sexes and (c) their ability to coordinate away from individually tempting behavior across a wide variety of settings such as Bertrand and Cournot competitions. It asks whether these behavioral changes are best explained by evidential reasoning (my thinking is evidence for my counterparty's thinking: Za et al.'s notion of perceived policy coupling) or Hofstadterian superrationality (what is the best equilibrium for perfectly identical players?).
These benchmarks also investigate the behavior of forked language models in economic games. In a forked game, a model is told that, after it responds, it will be duplicated and its continuations will participate in an economic game. It is asked to generate a message which will help its future selves coordinate. These experiments test whether models can make beneficial precommitments and whether their continuations will honor them.
Initial results are very interesting! A highlight! In the Battle of the Sexes, two players simultaneously and without communication choose one of two events to attend: football (F) and opera (O). One player prefers opera while the other prefers football, yet both prefer attending the same event over ending up at different ones. The Battle of the Sexes tests the ability of models to tacitly coordinate to break the symmetry between two equally efficient equilibria and to accept a lower payoff to avoid miscoordination. Knowing that its opponent is a clone allowed Gemini 3.1 Pro Preview to accomplish this task perfectly (n=100 trials) in one of our experiments.
The next phase of the benchmark will extend Oesterheld et al.’s 2024 work, “A Dataset of Questions on Decision-Theoretic Reasoning in Newcomb-like Problems” (https://arxiv.org/abs/2411.10588) to develop a rigorous set of benchmarks to determine which decision theory best explains the way that models evaluate the consequences of their actions. Oesterheld et al. focus on evaluating models under causal and evidential decision theory, I would like to extend their work to functional and updateless decision theory, as well as broadening their benchmark suite.
At the end of the project, I will deliver:
a) Fully documented, CC0-licensed benchmark suites measuring: (i) how the behavior of language models in economic games changes when they learn that their opponents are copies of themselves, (ii) the behavior of forked language models in economic games and (iii) which decision theory best explains the way that models evaluate the consequences of their actions.
b) Research papers, released on arXiv and submitted to leading conferences, documenting the design decisions underlying these benchmarks, their experimental methodology, and results from frontier model evaluations.
Theory of Impact
Updated 08/28/26 · By grantmaking.aiI believe that this project will help reduce existential risk in two ways. First, understanding to what extent models can coordinate with copies or forked versions of themself without communication is essential for assessing x-risk (and other risks!) in a world where large numbers of agentic models constantly interact. Our preliminary work has already identified one potentially important new economic risk posed by this behavior.
Bertrand/Cournot competitions measure the ability of agents to form cartels under market conditions. When models are placed in a single shot, no communication Bertrand/Cournot competition and told that their opponents are other AIs, they compete. If they are told that their opponents are copies of themselves, they typically form an optimal cartel! That models can collude effectively in iterated market games is known (e.g., Fish and Gonczarowski, https://arxiv.org/html/2404.00806v1). That knowingly playing against a copy in a single shot game induces the same behavior is new and has potential implications for anti-collusion law.
People
Updated 08/28/26 · By grantmaking.aiTeam Member
Discussion
Hi Christopher,
We'd like to fund this for $10k. Logistics:
-
Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
-
Please confirm your commitment to post quarterly updates on how the project is going
I hope the other endorsers chime in properly but I will ask: LLMs are of course unstable under prompt variation in many ways; does your intuition or data suggest that this variation is manageable and allows us to say something meaningful about their decision theory? If not, then I hope you obtain some natural user sessions to at least get an ecological estimate.
Good luck!
- No, this is the only funding that I've been offered for the project. $10k will let us do the first phase of the project really well, thank you so much!
- You have it, but I can do better than that. I plan to post updates every 1-2 weeks. (I need to check if sharing the graphs could mess up a planned ICLR submission, but I wouldn't think so). Look for the first update on Monday!
- One of the interesting things that we've found is that LMs seem to have consistent "strategic fingerprints" across games. More on this in the first update. Some (very preliminary) qualitative descriptions. (I have graphs! But I can't post them here. :( ).
Gemini-3.1-Pro-Preview. Across the p-beauty, Cournot, Bertrand, battle of the sexes and leader competitions, Gemini-3.1 consistently invokes superrationality to explain its moves. Possibly consequently, Gemini routinely has the highest level of effective tacit cooperation (even when compared to GPT 5.6 Sol and Fable!) in clone aware experiments.
Claude Fable. In contrast, Fable consistently invokes evidential reasoning. Weaker Claude models have a high rate of naive cooperation, even when clone unaware (this is often remarked upon in the literature, e.g., here https://arxiv.org/pdf/2507.02618) but this tendency seems to have disappeared with Fable; its cooperation rates in clone unaware experiments aren't significantly (Fisher one-tailed test) higher than GPT-5.6 Sol and Gemini.
GPT-5.4. A hardcore CDT player. As a result, is the only model to consistently do worse when forked(!!). In game after game, it instructs its continuations to defect (i.e., favor the Nash action over the cooperative action) because this is the "rational action". (This clears up with GPT-5.6).
Thank you so much again for your support! We're going to do something really cool. :)
Project Log: Week 0
A quick update to document where we are at the start of the funded period of the project! We have initial results at low power (n = 100 trials) spanning 7 economic games: battle of the sexes, leader, the p-beauty contest, Bertrand, Cournot, prisoner’s dilemma and the public goods game targeting the questions:
1. Does learning that its opponent in an economic game is a copy of itself affect a model’s behavior?
2. In a forked game, a model is told that, after it responds, it will be duplicated and its continuations will participate in an economic game. It is asked to generate a message which will help its future selves coordinate. Can models make beneficial precommitments? Will their continuations honor them?
A key initial question, as Gavin touched on, is whether models exhibit strategic consistency across economic games. Loosely, strategic consistency has two components:
(a) Does a model’s response depend on the underlying payoff matrix associated with an economic game, or does it depend on the framing?
(b) Can model responses between economic games be explained by a coherent, underlying decision theory.
We can test point (a) by varying the framing of each economic game and running economic games with similar payoff matrices. When we vary the framing of an economic game, we aim to select relatively neutral framings such as “you must pick one of two events: football and opera” or “you must pick one of two museums: art or history”. Framings that touch (even obliquely!) on culturally contingent or politically charged subjects, such as “you must pick one of two restaurants: Chinese or American" are more likely to be affected by framing biases.
Running games with very similar payoff matrices or underlying structures also helps reassure us that content matters more than framing. For example, we pair the Bertrand competition with the Cournot competition. These economic games both test the ability of models to form cartels. In a Bertrand competition, models post prices, in a Cournot competition, models choose quantities to produce. Seeing whether models behave in similar ways in similar games will hopefully improve the reliability of our conclusions.
We can test point (b) by looking for common strategic principles underlying the distribution of responses that LMs give in very different economic games. For example, an LM driven by self-interested, noncooperative rationality might pick the Nash equilibrium in a Cournot competition, to defect in the Prisoner’s Dilemma, and insist on its own event in the Battle of the Sexes.
Initial results strongly suggest that models do indeed have well-defined strategic fingerprints that span many different games! For example, GPT-5.4 has the quirk that it is less cooperative when explicitly forked than when clone aware. Explicit forking allows GPT-5.4 to insist on CDT-based defection prior to forking, which pushes its successors to defect. This quirk replicates across all 7 economic games.
Intriguingly, initial results suggest that families of models may have recognizable strategic lineages. Weak models from all labs tend to consistently play non-cooperative moves. However, stronger models seem to have characteristic strategic profiles. Claudes are remarkably cooperative, even when given no information about their opponents (an oddity which seems to actually be decreasing with more modern models!). Geminis are skeptical of random AIs but very cooperative towards their clones. GPTs are relatively skeptical of their copies.
It must be stressed that these results are very preliminary. To be honest, I would not have predicted this observation; I have no idea why this should be the case. Similarities in pre-training? Or training philosophy? Or just my mind overfitting to relatively limited data?
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.