grantmaking.ai Launch Round
There are three high-priority categories that could be covered by this benchmark: nuclear weapons, biological threats, and loss of infrastructure. The minimum vs ideal amount mainly differs in how many categories we cover, in what depth, and with which experts.
The minimum amount would allow us to create the benchmark for one category using an expert consensus from our existing and internal networks. This funding would cover 1 FTE-equivalent for 6 months, overhead, and modest funding (~$5k) for subcontracting help with evals and/or tokens. ALLFED has done extensive research on meeting food and water needs following a nuclear war and could create an eval with relatively little external input. We also have multiple projects on protecting the public and meeting basic needs in biological catastrophes and loss of infrastructure scenarios, and could alternatively start with an eval of one of these categories.
The maximum amount would allow us to set up dedicated expert working groups to create benchmarks for all three categories. This would cover 2 FTE-equivalents for 12 months, overhead, a higher subcontracting and tokens budget (~$30k), a publication, and small stipends for field experts, which would be particularly beneficial for biological threats and loss of infrastructure scenarios. This is similar in costs to WMDP Benchmark, which states in its preprint that it cost over $200k and covers biosecurity, chemical security, and cybersecurity from an offensive perspective (https://arxiv.org/pdf/2403.03218). For any amount within the funding range, we would publicize the results via the AI Security Institute database and a website with a leaderboard of model results.