grantmaking.ai Launch Round
$5,000 (minimum): rent the GPU hours for one complete RL training run of a small open code model against the verifier reward, audit every checkpoint against the held-out signals, and publish the divergence curve and exploit taxonomy, flattering or not. The environment and audit set are already built, so this is compute plus my time.
$25,000 (ideal): the same experiment done properly. That means three model scales, multiple seeds, and the ablation that matters, which is training with and without the fail-closed gate to see whether fail-closed verification suppresses gaming or just hides it. RL is hundreds of runs, so the real constraint is iteration time, and rented GPUs meter it since every run costs money and I ration them. If the workstation in my other application is funded, these iterations cost electricity instead and this money shifts to more scales and seeds. Plus a published dataset of every exploit found, so other groups can study the failure mode without rebuilding the rig.