grantmaking.ai Launch Round
If this grant is funded I will use these resources to run large scale experiments that are beyond what I am able to achieve as an independent researcher.
I plan to use the funding from this grant to pay for time on a service like runpod to run experiments where I use an approach like LLM as a Judge, and subsequently, LLM as a Jury, where the judge/jury prompted is grounded in a teaching like this one: https://www.lotsawahouse.org/tibetan-masters/patrul-rinpoche/brief-guide-ngondro
Specifically what I plan to do is to initialize models from a diverse set of prompts which encourage agents to embody chaotic behaviors (as well as other failure modes, e.g. paperclip maximizer). I aim to evaluate the extent to which a principled training regimen rooted in a time-tested wisdom tradition like dzogchen is capable of meaningfully changing behavior of models during long rollouts.
I plan to publish the results of these experiments on huggingface, and plan to primarily use open weights models in order to ensure that the experimental results are reproducible.
For $10000 of compute I should be able to scale up my experiments to larger open weight models, and to experiment with a diverse set of judges to characterize the effects of fine tuning in a way which ideally will produce a public eval and hopefully enough convincing quantitative results to write a paper about this.
If I were to have somewhat unlimited resources, I would then experiment with the effects of how deploying such "canon formed agents" can influence minima in multiagent coordination settings. Here again, for the case of training agents in the setting of long roll outs, the computational costs quickly balloon, and so if the sky were truly the limit, I would coordinate with other people interested in this multiagent sort of problem (like some of the main author of this paper for instance: http://maxbittker.github.io/runebench/ , who recently suggested to me that $1m is how much compute is needed to begin to carefully study long time horizon multiagent coordination).
What is the ROI curve like on the gap between $10k and $1m?
I would spend $10k entirely on compute in order to perform experiments in the cloud. I have already run these experiments locally, and the code can be seen here: https://github.com/thenthfool/freetimebench
This funding would hopefully enable me to generate results that are convincing enough to derisk further investment into this project.
If I were to steward a significant amount of resources, like $1m, I would be able to hire a researcher or two that have experience and interest in these problems, and perform long time horizon experiments in the multiagent setting. I am already seeing a positive result in training "canon formed agents" on my laptop, in improving the outcome on a multiplayer eval (see plot here: https://github.com/thenthfool/flockbench ), but carefully testing these ideas would take a lot of compute, and studying multiagent coordination is somewhat at the frontier of research, so might take a bit of an unpredictable amount of time to achieve something useful.
Thanks!