grantmaking.ai Launch Round
Understanding how scaling inference computations affects model performance is critical for predicting the performance of future frontier models. Understanding whether there is a log-linear relationship between computation and performance for agents will help us more accurately predict the pace of R&D automation and other risks associated with autonomous systems. Furthermore, most modern EVALs make little use of harnessing, resulting in somewhat underestimated estimates of model performance on real-world tasks, and we are interested in quantifying this bias.
You can check our prototype!
We've already completed significant preliminary work!
- We analyzed existing research on scaling laws and agent problem evaluation methods.
- We developed an evaluation pipeline using Inspect/Petri that allows for varying the level of computation.
- Using this reproducible infrastructure, we obtained scaling law graphs for a wide range of agent problems using the GAIA benchmark, which can be found in this repository: https://github.com/gornkv/gaia_pipeline
However, at the current cost of runs, we still need to invest in computations to improve representativeness.
On a Tesla V100, which is sufficient to run all the required models, the rental cost is $0.70/hour. Running level 1 and level 2 GAIA tasks on it, sufficient for statistical significance, takes approximately 20 hours during optimization. Therefore, to create a baseline of 14 techniques in raw form, we need a total of $0.70/hour x 20 hours x 14 techniques = $196.
Based on baseline analysis and subsequent evaluation of statistically significant 16 combinations of these techniques (3-4 techniques at a time), we get $0.70/hour x 20 hours x 3.5 techniques x 16 combinations = $784 extra, for a total of $196 + $784 = $980 per model of the qwen 3.5 - 9b type.
Since we're interested in at least six model types for reproducibility, we charge ~$6,000