A video game that teaches people about AI alignment and race dynamics.
A video game that teaches people about AI alignment and race dynamics.
Project Details
Updated 07/12/26 · Provided via application · VerifiedThe Alignment Game (alignmentgame.net, title subject to change and that link is a beta version) is a medium fidelity interactive simulation wherein the player controls a frontier AI lab. Their objective is to reach aligned ASI before any rivals reach ASI and before any models cause catastrophes. The thesis of the game (what a player would be convinced of after playing) is that alignment is not necessarily easy, that race conditions significantly narrow humanity's margin for error, and that strong governance and well-enforced regulations widen that margin. The game is designed to make misalignment feel foggy during an unskilled playthrough but intelligible in retrospect, and leave a non-technical player with an understanding of eval awareness, safety research, regulation, and misalignment. At the end of the game, we would point players to further reading and actionable steps (vote, write to your congressperson, and consider working in AI safety, with some pitch that lots of people can be useful to the movement and very few are) + a shareable endscreen.
Gameplay is turn based, with turns lasting one quarter. Each turn, players and rivals decide where to allocate spending and "work budget" between capabilities research, safety research, compute, lobbying, and litigation. Capabilities and safety research items are described to the player, along with diegetic warnings about potential risks and failure points (that link to real research, for those who want to learn more). Players must balance racing ahead to maintain market dominance without cutting too many corners and producing a misaligned superintelligence. This balancing act plays out in many decisions. Players decide what research items to research and the extent to which to utilize their own AI models to accelerate development across each research item they unlock. Players who use no AI assistance fall behind, but players who use it too heavily, especially on safety items, are likely to end with a misaligned, extremely eval-aware model. In particular, players who use significant AI assistance to carry out mechanistic interpretability research or other evaluations will find that their models appear extremely aligned, regardless of the model's true dispositions. Investment in safety early on is crucial, since as the model's capabilities grow, eval awareness makes it increasingly costly to get reliable safety indicators, and more difficult to reduce misalignment.
Players and rivals can lobby for or against certain regulatory policies, engage in court fights about enforcement, and comply with regulations or risk penalties. Their effectiveness in lobbying is determined by their valuation and their lobbying spend, the general regulatory appetite (which grows as AI models become more dangerous), and the lobbying spends of rivals. These factors also help determine the initial strength of enforcement. Players who make good use of lobbying early on can slow down the race, punish reckless labs, and widen their margin for error. Players who are too slow to push regulations through might get bogged down in court fights, pausing enforcement as the race takes off, and those who ignore regulation altogether are likely to lose to reckless rival labs. On an initial playthrough, players are likely to make certain mistakes (eg researching chain of thought reasoning before expanding pre-training datasets) that will slow their progress and force them to behave somewhat recklessly (high AI assistance, little expensive safety research, release despite warning signs) or lose market share. In the process, they are likely to lose (at least on "realistic" difficulty), either by losing the race to another company, or by producing a misaligned ASI. The end screen gives a post-mortem, which shows how key decisions affected the underlying model, resimulates with different individual choices to give counterfactual advice, and provides tailored tips (eg "You carried out mechanistic interpretability research with full AI assist. This allowed the model to hide its underlying dispositions from you" or "You didn't spend anything on lobbying. Strong regulation is a key mechanism to slow the race and let safe labs win out.") On subsequent playthroughs, players can work through easier difficulties, incorporate feedback, engage with more aspects of the game, and eventually win in realistic mode.
I've made a prototype. I'll continue refining the gameplay, but to make a game that could go viral and reach a large audience, I'll need to pay a composer/arranger to make music (or search for a volunteer), spend on marketing, and scale the infrastructure to serve a larger audience.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiTL;DR: Like AI 2027, but shown and felt rather than told. Pays off very well if one person is convinced to make a career pivot towards AI safety work, and that is achievable with a very low (0.001%) conversion rate from playing some of the game to pivoting careers. At that conversion rate, 100k would get new people working in AI safety at a rate such that, if scaled, could double the size of the AI safety workforce (~1000 people) for roughly 100 million dollars (not suggesting this would scale that far, but that if we had an opportunity to double the workforce for 100 million dollars, we would, and so we ought to fund this project).
I have a lot of smart friends, many in STEM or CS-adjacent areas, with very bad mental models of alignment and catastrophic risks from AI. I've found that narrative explorations of AI risk effectively explain why we can't just "program the model to be good" and how misalignment enters a model, but often fail when people ask why we can't just slow down and behave safely. This game is designed not only to teach them where misalignment originates and why it is difficult to root out, but also why the decision-making environment that frontier labs work in is dangerous and scary, and how regulation might help. If the game succeeds in reaching those uninformed users, it could convince many of them that they ought to vote with AI safety in mind, and some that they can and should work in governance or technical safety research.
Suppose, with a budget of $100k, that we spend $15k on development and then $85k on advertising. With a reasonable estimate of $1 to acquire a user who plays at least some of the game (slightly lower than a typical cost per install of a mobile android game. If this seems low, consider that the game is free and requires no installation), that turns out to 85,000 users. To add, in expectation, one new person involved in AI safety research or policy work, we would need roughly a 0.001% chance of convincing any individual player to pivot (10% play to the end, of them 10% of users can usefully contribute, of them 0.1% proceed to make a career pivot). With well-targeted ad campaigns and careful messaging in the game, that is more than plausible. The tail of the value distribution also goes out quite far. Virality compounds, and with strategic marketing and some luck, we could well reach into the hundreds of thousands of users (as a simpler video game about AI alignment, Universal Paperclips, did in 2017, when the discourse was much more niche).
Targeted ads are a ceiling on our user acquisition costs. With some luck on YouTube partnerships with value-aligned creators (, ), we could end up with much lower user acquisition costs than $1. Our ability to make these larger commitments depends on the amount of funding received, but more money likely means we will be able to reduce acquisition costs for a good while before diminishing returns set in.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Funding Details
- May 20, 2026
- -
- 3-4 months of development
- -
- -
- -
- -
- -
- -
- -
Discussion
No comments yet. Be the first to share your thoughts.