grantmaking.ai Launch Round
Minimum version — $30,000 over approximately four months
The minimum version funds one lead researcher, compute and activation storage, external technical review, and reproducible publication of the results. It will use one small open-weight model and a limited set of core experimental conditions: short versus long horizons, exclusive versus shareable control, and enforceable versus unenforceable commitments.
The minimum outputs will be:
- a reproducible three-agent control-allocation environment;
- a behavioral dataset and phase diagram;
- baseline behavioral and activation-based detectors;
- targeted circuit analysis of selected coalition-formation or betrayal decisions;
- causal activation interventions where technically feasible;
- an open-source code release and research report.
Ideal version — $50,000 over approximately six months
The ideal version adds a part-time research engineer, substantially more compute, additional strategic conditions, a second model or architecture, more saved RL checkpoints, stronger out-of-distribution evaluation, and fuller causal interpretability work.
The additional funding would support:
- replication across model families or sizes;
- analysis of how strategic representations emerge during training;
- more comprehensive cross-layer transcoder or sparse-feature analysis;
- larger held-out evaluation sets;
- expert methodological review;
- polished documentation, datasets, and publication-quality results.
The project will use stage gates. Behavioral experiments and lightweight probes will first identify specific decisions worth analyzing. Expensive circuit-tracing work will be applied only to selected, reproducible strategic transitions rather than indiscriminately across all episodes.