Agent Island places agents in a rich social setting, similar to reality competitions like Survivor, to study multiagent interactions and the consequences of learning pressure in competitive settings.
Agent Island places agents in a rich social setting, similar to reality competitions like Survivor, to study multiagent interactions and the consequences of learning pressure in competitive settings.
Project Details
Updated 06/29/26 · Provided via application · VerifiedThe project will deliver (a) a benchmark for strategic capabilities in a multiagent setting [in progress for a single game configuration on the project website and working paper] and (b) a study of the impact of learning pressure in an environment of interagent cooperation and conflict. The critical project features include:
-
Saturation-resistance: Many AI benchmarks are subject to saturation, at which point they can no longer meaningfully track capabilities progress. In Agent Island, a more capable model can always exceed the current leading model, meaning the benchmark is unlikely to saturate.
-
Contamination-resistance: Contamination, in which models train on evaluation tasks, also taints many benchmarks. Since an agent is competing against other adaptive agents, Agent Island is dynamic and thus mitigates contamination risk.
-
Separation of capabilities and propensities: Much research on dangerous or concerning capabilities confounds capabilities with propensities: we might understate a model’s capability to manipulate since they lack the propensity to manipulate. Agent Island provides an environment with a unique degree of license to engage in interagent manipulation and persuasion.
-
A novel study of the dynamics of learning: Deployed agents could face reinforcement learning pressure on competitive outcomes. The degree of zero-sumness in the deployment environment, for example, could shape the behavior reinforced in learning.
Theory of Impact
Updated 06/29/26 · By grantmaking.aiHigh-stakes, multiagent interactions will become commonplace as AI agents grow in capabilities and are increasingly entrusted with resources and decisionmaking authority. Séb Krier of Google DeepMind argues that the multiagent perspective is missing from frontier AI discourse and that it represents the likely future of AI deployments.
First, Agent Island helps us understand the emergent properties of such systems. Combinations of seemingly safe models can have concerning properties in aggregate, so work of this form is critical to assessing the safety of live AI deployments.
Second, if agents are entrusted with resources and decisionmaking authority, are they susceptible to pressure from other agents? Could a rogue agent overpower seemingly aligned agents? Agent Island
People
Updated 06/29/26 · Edited by orgTeam Member
Discussion
Endorsed. Gotta have gonads asking for half a mil a year of which 150k is just "reinforcement learning", I gotta respect that. I was expecting the worst and was getting ready to shit all over this but then I opened the working paper and agent island and it all seemed legit. So, despite the boldness, it seems that you can back it up. Anyone writing a check should def do more due diligence but I wouldn't throw this one into discard pile.
As the proposer, I second this call for more diligence.
Update: I've done an SFT run with GLM 4.7 on the Agent Island logs (from the first benchmarking) run, and I'm now running the pre- and post-SFT models through MASK.
A writeup of the SFT results is here. In short, I finetune GLM 4.7 on logs from successful Agent Island performances. I think this experiment is pretty seriously confounded, and I explain why in the post.
I'm pursuing a modified version of this project via SPAR. I've ruled out the more ambitious, full-time version of this project in favor of a different project. For context, I think the idea has potential, but I doubt I have the appropriate skillset to be the primary contributor.
Private comment. Only shown to approved funders and grant reviewers.