grantmaking.ai Launch Round
The Alignment Game (alignmentgame.net, title subject to change and that link is a beta version) is a medium fidelity interactive simulation wherein the player controls a frontier AI lab. Their objective is to reach aligned ASI before any rivals reach ASI and before any models cause catastrophes. The thesis of the game (what a player would be convinced of after playing) is that alignment is not necessarily easy, that race conditions significantly narrow humanity's margin for error, and that strong governance and well-enforced regulations widen that margin. The game is designed to make misalignment feel foggy during an unskilled playthrough but intelligible in retrospect, and leave a non-technical player with an understanding of eval awareness, safety research, regulation, and misalignment. At the end of the game, we would point players to further reading and actionable steps (vote, write to your congressperson, and consider working in AI safety, with some pitch that lots of people can be useful to the movement and very few are) + a shareable endscreen.
Gameplay is turn based, with turns lasting one quarter. Each turn, players and rivals decide where to allocate spending and "work budget" between capabilities research, safety research, compute, lobbying, and litigation. Capabilities and safety research items are described to the player, along with diegetic warnings about potential risks and failure points (that link to real research, for those who want to learn more). Players must balance racing ahead to maintain market dominance without cutting too many corners and producing a misaligned superintelligence. This balancing act plays out in many decisions. Players decide what research items to research and the extent to which to utilize their own AI models to accelerate development across each research item they unlock. Players who use no AI assistance fall behind, but players who use it too heavily, especially on safety items, are likely to end with a misaligned, extremely eval-aware model. In particular, players who use significant AI assistance to carry out mechanistic interpretability research or other evaluations will find that their models appear extremely aligned, regardless of the model's true dispositions. Investment in safety early on is crucial, since as the model's capabilities grow, eval awareness makes it increasingly costly to get reliable safety indicators, and more difficult to reduce misalignment.
Players and rivals can lobby for or against certain regulatory policies, engage in court fights about enforcement, and comply with regulations or risk penalties. Their effectiveness in lobbying is determined by their valuation and their lobbying spend, the general regulatory appetite (which grows as AI models become more dangerous), and the lobbying spends of rivals. These factors also help determine the initial strength of enforcement. Players who make good use of lobbying early on can slow down the race, punish reckless labs, and widen their margin for error. Players who are too slow to push regulations through might get bogged down in court fights, pausing enforcement as the race takes off, and those who ignore regulation altogether are likely to lose to reckless rival labs. On an initial playthrough, players are likely to make certain mistakes (eg researching chain of thought reasoning before expanding pre-training datasets) that will slow their progress and force them to behave somewhat recklessly (high AI assistance, little expensive safety research, release despite warning signs) or lose market share. In the process, they are likely to lose (at least on "realistic" difficulty), either by losing the race to another company, or by producing a misaligned ASI. The end screen gives a post-mortem, which shows how key decisions affected the underlying model, resimulates with different individual choices to give counterfactual advice, and provides tailored tips (eg "You carried out mechanistic interpretability research with full AI assist. This allowed the model to hide its underlying dispositions from you" or "You didn't spend anything on lobbying. Strong regulation is a key mechanism to slow the race and let safe labs win out.") On subsequent playthroughs, players can work through easier difficulties, incorporate feedback, engage with more aspects of the game, and eventually win in realistic mode.
I've made a prototype. I'll continue refining the gameplay, but to make a game that could go viral and reach a large audience, I'll need to pay a composer/arranger to make music (or search for a volunteer), spend on marketing, and scale the infrastructure to serve a larger audience.
Minimum case:
Base costs:
- 2500 for music
- ~250/year of web infrastructure (depends on reach), call it 250 from this initial grant
- 2500 for design contracting and story consulting
- 3000 to pay myself (my personal expenses are low, but this allows me to prioritize the project)
- 3000 for ad creatives and production
Fixed costs total to $11,250. This leaves $18,750 for marketing, which would mostly be targeted YouTube or other ads.
Ideal case:
Base costs:
- Same $11,250 fixed costs
Marketing ($88,750):
- At this scale, we could partner with multiple large YouTubers with educated, technical audiences. I am waiting on a quote from Kurzgesagt, and will update this when I hear back.