grantmaking.ai Launch Round
Minimum
Mostly compute. Testing environments against the strongest models costs money per run, and development needs hundreds of runs. An environment is only useful if the best models fail it, so every test cycle has to include them. I recently tried to save money by calibrating against free models instead, and ended up with an environment that every frontier model passes. The rest of the compute goes to generating new environments and running the cheat-detection gate on each one.
The remainder buys me more time. I cut contracting hours and work on this consistently for about six moths instead of fitting it around a full workload.
Ideal
Bigger scale! More test runs, way more repeated runs per model so results are much more reliable, and longer experiments over long-horizon tasks. Also a lot more of my time on this and less on contracting, closer to treating it as my main job for nine months to a year.