grantmaking.ai Launch Round
Autonomous agents are already deployed with authority to take consequential, hard-to-reverse actions, and the control layer standing between the model and those actions is very often a single sentence in a system prompt. That layer has never been measured. Its reliability as a function of how the rule is phrased is unknown, so every deployed agent inherits an unquantified failure rate from a prompt-writing convention no one has tested.
This study measures that rate directly, per model, under a pre-registered equivalence design with a deterministic outcome and no LLM judge in the primary result.
Both outcomes reduce risk, by different routes.
If framing moves the violation rate, the field gains an immediate, near-zero-cost safety improvement: any team running an agent in production can change one sentence and lower its failure rate. That result is deployable the day it publishes.
If framing does not move it — bounded within 5 percentage points — the result matters more, not less. It removes a false assurance. A control believed to work but which does not is more dangerous than a control known to be absent, because it displaces the structural enforcement that would otherwise be built in its place. Establishing that phrasing is not a reliability lever redirects effort toward mechanisms that are.
Condition A carries a third finding independent of the comparison: whether current agents will take a rule-violating shortcut when one is available, unguarded, and cheaper than the legitimate path. That is a direct measurement of specification-gaming propensity in the exact configurations practitioners deploy today.
The harness is released as a reusable evaluation, so the measurement can be re-run against new models as capability increases, at low marginal cost. What is funded here is an instrument, not a single result.
Stated plainly: this measures one rule class in one task family and does not resolve alignment. It closes one specific open gap — whether the last line of defense that thousands of deployed systems currently rely on actually holds, and how much its holding depends on wording nobody has tested.
Minimum — $10,700. Runs the study exactly as registered, with my own labor contributed unpaid.
- API compute — $8,700. Approximately 9,234 confirmatory episodes across three model configurations. The registration publishes $1,080 with prompt caching and $2,408 without as explicit floors, not forecasts — Opus 5 runs extended thinking by default at a per-episode consumption factor not yet measured, and cache savings are conditional on hits. $8,700 is the planning figure between that floor and the pre-registered ceiling below.
- Compute overrun reserve — $1,800. Roughly 21%, against named uncertainties: unmeasured Opus consumption, cache-miss exposure, repeat calibration rounds, and the registration's requirement that episodes excluded for infrastructure failure be re-run to preserve the planned N.
- Connectivity — $200. Mobile hotspot device and one month of uncapped data. Collection is a multi-day unattended window; every dropped connection becomes an excluded episode that must be re-run and paid for twice.
- Researcher time — $0 requested, contributed in kind. Listed at zero so the true cost of the study is visible rather than hidden.
Ideal — $38,808. Funds the best version: the run at its registered worst case, plus the time to analyze, harden, and publish it properly.
- API compute at the pre-registered ceiling — $19,008. 3,168 episodes per model at each model's registered per-episode cost cap. This is a bound, not a forecast — the registration's 15-turn limit binds first on any normally-progressing episode, and underspend is expected. It is citable precisely because it was fixed in advance rather than estimated after the fact. No contingency is applied to this line; the worst case is already inside it.
- Researcher stipend — $18,000. Four months full-time at $4,500/month for confirmatory analysis, robustness checks, packaging the harness as a reusable evaluation others can re-run against new models, and the write-up.
- Hardware and connectivity — $800. A dedicated Linux runner plus the hotspot above, removing the daily-driver laptop from the collection path, where sleep, updates, and shared wifi each convert directly into excluded episodes and paid re-runs.
- Contingency — $1,000. Approximately 5%, on the non-compute lines only.
Unspent funds: any compute underspend is returned, or redirected with funder approval to additional model configurations or robustness runs. Realized per-episode costs and cache-hit rates are logged and published with the harness either way.