grantmaking.ai Launch Round
Any funding I receive will go towards a six month project where I explore the likely utility function of advanced artificial intelligences using concepts and techniques from evolutionary game theory. In particular, more funding allows me to:
- focus on this project full-time
- spend time designing more sophisticated and informative simulations
- purchase server time to run these simulations
The output would be the results of the simulations, in addition to my write up. I plan to publish all results and research online.
Here are a few claims I take it are widely endorsed in the AI safety literature.
- It is very difficult to glean information about the utility function a superintelligence would have
- It is also fairly difficult to glean information about the utility function a less than super intelligent AI would have
- We can infer some things about the approximate goals both of less than super intelligent and superintelligent ais
- What we can glean is at the very least worrying for humanity
- Any further information on any of these topics is highly valuable
This is where evolutionary game theory comes in. Although selective pressures and evolution are concepts that have been lurking in the AI safety literature for some time (Bostrom 2014, Carlsmith 2022, Kelters 2025) they have rarely been given direct focus and almost universally not in the technical way someone trained in evolutionary game theory would approach them. This is a gap in our understanding of advanced artificial intelligence that it is vital to fill.
In my research so far I've already identified some results which are counterintuitive. The central highlight is that there appears to be reason to think that that advanced artificial intelligences with rational pure time discount rates would be systematically outcompeted by those with irrational pure time discount rates. I've also found interesting results relating to EDT and others to do with the self-sampling assumption.
All knowledge is provisional and on further investigation it might turn out these results are not robust enough to be interesting; features of my setups that now seem plausible or inessential to the result may turn out to be neither. Nonetheless, such research is valuable insofar as we consider a greater understanding of the likely motivations and behavior of advanced artificial intelligences valuable.
I take it that for research to have a useful impact two things must happen: the research must generate new and useful information, and that information must be disseminated. Any funding provided would assist with both of these tasks.
I take it the information that would be generated is useful because the more we understand about the motivations and behavior of near future artificial intelligences the better we are able to design safety procedures that will a) prevent catastrophes happening or that will b) help us recover from small catastrophes/prevent them spiraling into large ones.
Lots of existing research has been done into what decision theories or decision theoretic principles existing large language models show a preference towards (e.g. in the Claude Sonnet 5 System Card). If we enter a regime where current or near future artificial intelligence agents are subject to intense selective pressures with each other it could be that they would come to adopt different decision theoretic principles than one would predict by observing them currently. This has wide implications, such as for the design of boxing methods and potential proposals for trade with artificial intelligences (e.g. in Oesterheld 2017, also referenced in AI 2040: Plan A (2026)).
In particular, if we could predict that advanced artificial intelligences would likely possess extremely high pure time discount rates, they would plausibly not try and escape a box they were put in. With a zero pure time discount rate then a 1% chance of colonizing the lightcone is a very strong incentive (for a risk neutral agent). With even a moderate pure time discount rate, that incentive may be radically reduced, leaving agents with much less incentive to escape the box.
This funding would also allow me to dedicate more time to sharing widely the results I do find, both in person and online.
Salary: 13.5k
Travel: 500-1000
I'd like to connect with other AI safety researchers, researchers in evolutionary game theory, and share any interesting results widely. My current plan is to travel to LSE in London, LMU in Munich, and the Bay Area. Meeting other researchers in those two disciplines is vital for helping the project go as well as possible, sharing the results widely is of course a value multiplier.
Software: 1000
I can code, but I'd like to pay for 4 months of Claude 20x ($800) and 2 months of Claude 5x ($200). I've found it best for my coding needs. Historically I've largely used it for prototyping and UI design, and I'd expect the same here. Both of these use cases make my research process much more efficient. (But I wouldn't necessarily need the higher tier for the whole project, hence the split.)
Books: 250-400
My university library is excellent, but there's been a lot of times I've had to purchase books myself.
Server time: 150-250
For running the simulations. I do not have a home setup for running large projects like this and it would be inefficient to create one when the usage will be intermittent throughout the project.
Contingency (not included in minimum): 1000
Min: 13,500+500+1000+250+150=15,400
Max: 13,500+1000+1000+400+250+1000=17,150