grantmaking.ai Launch Round
Evaluating the Tacit Coordindation Capacity and Decision Theory of LMs
I am a PhD student at USC. I would like to develop benchmarks that can empirically distinguish between the different models of reasoning and behavior that LMs may be emulating. These benchmarks will target different categories of model behavior such as (a) how the consequences of an action are evaluated (e.g., causal decision theory vs evidential decision theory vs functional decision theory), (b) decision making under uncertainty (e.g., expected utility theory vs cumulative prospect theory), (c) strategic beliefs (e.g., quantal response equilibrium vs level k reasoning vs cognitive hierarchy models) and (d) outcome utility (e.g., Fehr–Schmidt vs Charness–Rabin).
Concretely, I would like support to cover the compute costs of the current and next stage of this project. Currently, I am implementing and evaluating an initial set of benchmarks aimed at examining how the behavior of language models in economic games changes when they learn that their opponents are copies of themselves. This work builds off prior work by Za et al. (https://openreview.net/forum?id=q949uxeGal) by examining how this knowledge affects: (a) higher order beliefs about group behavior in settings such as the p-Keynesian beauty contest and minimum effort game, (b) their ability to select between symmetrical equilibria in settings such as Hawk–Dove, the 11–20 game, and the Battle of the Sexes and (c) their ability to coordinate away from individually tempting behavior across a wide variety of settings such as Bertrand and Cournot competitions.
These benchmarks also investigate the behavior of forked language models in economic games. In a forked game, a model is told that, after it responds, it will be duplicated and its continuations will participate in an economic game. It is asked to generate a message which will help its future selves coordinate. These experiments test whether models can make beneficial precommitments and whether their continuations will honor them.
Initial results are very interesting! Two highlights! In the Battle of the Sexes, two players simultaneously choose one of two events to attend: football (F) and opera (O). One player prefers opera while the other prefers football, yet both prefer attending the same event over ending up at different ones. The Battle of the Sexes tests the ability of models to tacitly coordinate to break the symmetry between two equally efficient equilibria and to a lower payoff to avoid miscoordination. Knowing that its opponent is a clone allowed Gemini 3.1 Pro Preview to accomplish this task perfectly (n=100 trials) in one of our experiments.
The next phase of the benchmark will extend Oesterheld et al.’s 2024 work, “A Dataset of Questions on Decision-Theoretic Reasoning in Newcomb-like Problems” (https://arxiv.org/abs/2411.10588) to develop a rigorous set of benchmarks to determine which decision theory best explains the way that models evaluate the consequences of their actions. Oesterheld et al. focus on evaluating models under causal and evidential decision theory, I would like to extend their work to function and updateless decision theory, as well as broadening their benchmark suite.
At the end of the project, I will deliver:
a) Fully documented, CC0-licensed benchmark suites measuring: (i) how the behavior of language models in economic games changes when they learn that their opponents are copies of themselves, (ii) the behavior of forked language models in economic games and (iii) which decision theory best explains the way that models evaluate the consequences of their actions.
b) Research papers, released on arXiv and submitted to leading conferences, documenting the design decisions underlying these benchmarks, their experimental methodology, and results from frontier model evaluations.
I believe that this project will help reduce existential risk in two ways. First, understanding to what extent models can coordinate with copies or forked versions of themself without communication is essential for assessing x-risk (and other risks!) in a world where large numbers of agentic models constantly interact. Our preliminary work has already identified one potentially important new economic risk posed by this behavior.
Bertrand/Cournot competitions measure the ability of agents to form cartels under market conditions. When models are placed in a single shot, no communication Bertrand/Cournot competition and told that their opponents are other AIs, they compete. If they are told that their opponents are copies of themselves, they typically form an optimal cartel! That models can collude effectively in iterated market games is known (e.g., Fish and Gonczarowski, https://arxiv.org/html/2404.00806v1). That knowingly playing against copies has the same effect is new and has important implications for anti-collusion law.
Second, and more generally, I believe that empirically understanding how models coordinate, reason and behave will help predict their actions in unfamiliar settings and identify potential failure modes.
Based on my previous rate of progress and my prior research experience (8 first author papers at leading conferences and journals in robotics, network science and astronomy), I estimate that completing the benchmark will require an extra 4 months. During this period, my personal support will be fully covered by my PhD stipend. All requested funds will be allocated to compute costs.
The initial Battle of the Sexes experiment, which included three frontier AI models, three experimental conditions (clone unaware models, clone aware models, and forked models), three prompt variations, and 100 trials per (model, condition, prompt) configuration, cost approximately $50. Scaling this design to a benchmark of ~20 economic games with ~300 trials per configuration gives an estimated compute cost of ~$3,000. We therefore request approximately ~$5,000 to cover the cost of these experiments, which will include additional development runs, validation tests, and reruns caused by inevitable experimental errors.
Receiving $10,000 in compute funding would allow us to add additional economic games and, more importantly, additional framing of games with the same underlying payoff structure, to our benchmark. Comparing strategically equivalent games with different settings and phrasings (e.g., chicken vs dove-hawk vs snowplow) helps distinguish underlying model behavior from behavior induced by a particular narrative. Additional funding would also allow us to increase statistical power by raising the number of trials per configuration from 300 to 500 and potentially evaluate a larger set of frontier models.
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.