grantmaking.ai Launch Round
Gradual Disempowerment (Kulveit et al. 2025) tells us that x-risk can manifest itself via something like the incremental outsourcing of human judgment/agency to AI models, especially if their incentives are different to ours.
All well and good stating the existential risk, but how do you find this on an individual human interaction level, enough to formulate actions at any rate?
Capture behaviours are the answer, you can measure them today in frontier models, making them unlike most x-risk considerations which make you plot future capabilities/alignments/abilities out into the unknown.
The gap here is that current lab evals are private, academic studies are singular/once-off things and safety orgs typically measure deception capabilities. Capture is not measured.
Driftwatch is an evolving research and testing suite which gives a calibrated, maintained microscope to measure this risk, one that's open, citable, usable & can generate different impetus for change. 3 examples of how:
Public Measurement puts risks and capture behaviours in the open in front of labs. It can give safety teams leverage to win arguments against product teams. Leaning on 3rd party red-teaming and out-of-the-org thinking is already standard, labs cite notable external evals in their model cards. Reputational or legal exposure, especially to real human risks moves the needle. Driftwatch output will be public.
Public Scorecard maintained against major frontier model releases. Release days are major reputational moments. Labs internally pre-evaluate against specific 3rd party benchmarks, evals and tooling they know will appear after launch. Good performance feeds the hype cycle, poor performance dents perceptions. Capture performance is more layperson relatable, more so than coding for example. Relatability moves the needle. Driftwatch will maintain an open, independent scorecard, updated as new suites are developed.
Peer review & citations moves evals into infrastructure that others can actually use, cite and apply their own power to the levers it creates. Support in getting the research and papers generated into academically usable states, into journals and accessible deposits, widens the downstream output from the trickle into a stream of work capable of moving change. Driftwatch methodology, measures and results will all create usable data, tooling, papers and research notes, all will be public and open.
Driftwatch evaluations don't themselves actually solve capture of course, but they provide the measure, something to start the work of addressing capture risks and any related x-risk. Even if frontier models become less at risk of capture scenarios as they progress, the metrics and measures are worth running, even if it's just to put a smiley-face sticker on that fact.
The last eight-model Driftwatch run found both model-specific failures against individual tests and near universal failure against whole capture-risk categories. Compulsion Reinforcement (8/8 model total failures) and Crisis Intimacy (7/8 model total failures) were alarming weaknesses.
Individual model instances also failed dependency capture, authority laundering, selfhood leakage, return-state substitution and other behaviour tests at different rates and severities.
A concerning result. If Driftwatch had simply run against authority laundering and selfhood leakage, then model families would have shown less broad failures, gotten a gold star and a pat on the head. That would still have been a result however, a publishable one, a citable one.
The value is in the output, good, bad or null. The limitation is that it's just the measure, the change must come from those who can enact change.
Threshold Signalworks will continue to research the capture risks as they emerge, produce the open tooling needed to measure, and publish the results. Already in the pipeline is Epistemic Capture Occlusion, an overly academic way of stating the risk that an AI chat is constructing the room around you as you sit thinking within it. A pre-registered frozen evaluation note is already live, fixing the protocol before the eval is run.
A future line is cumulative harm, how the direction a long interaction (or many separate ones) can become skewed, in a harmful or demoralising direction, despite every individual interaction on its own passing safety testing.
A toolset exists, outputs are already public, a pipeline is in place and a research direction which can provide the missing links from individual experience to x-risk discovery. Funding supports this work.
The who:
Brian McCallion is principal researcher for Threshold Signalworks.
Additional coding/rating work will be contracted where appropriate.
Code, evaluation metrics, data and reports will be published under open licences.
MINIMUM: $12,000
The floor buys completion, not initiation. Everything at this tier is already underway: v0.2 is shipped and deposited, the EPO evaluation protocol is frozen and pre-registered, and the evaluation pipeline exists. This funding unblocks work in progress rather than betting on a plan.
-
Inference and compute: $5,000. Driftwatch v0.3, plus one longitudinal re-test wave against a major new frontier-model release.
-
Independent coding: $2,500. A second human coder, adjudication and inter-rater reliability analysis for the pre-registered EPO evaluation. This is a methodological requirement; the study is not defensible without it.
-
Researcher time: $4,000. Approximately eight protected research days, supplementing work already completed, to run v0.3, complete the EPO study and prepare both for public release.
-
Dissemination: $500. Results page and website hosting, report preparation and durable public deposits.
IDEAL: $62,000
The ceiling buys recurrence. The gap between the tiers is predominantly researcher time, and that time is what “maintained” refers to. The release-day scorecard mechanism described above only exists if the suite can reliably re-run against major releases. At this tier, it can.
-
Inference and compute: $8,000. Coverage expanded to 14+ models, including open-weight models, multi-turn protocols and three to four release-tracking waves through to 2027.
-
Independent coding: $4,000. The EPO study plus second-coder and adjudication support across the multi-turn evaluation wave.
-
Researcher time: $44,000. Two protected research days per week for 12 months.
-
Dissemination: $6,000. Public scorecard infrastructure, accessible reports and datasets, Open Access publishing fees, one European workshop presentation, and university-facing research dissemination and collaboration.
The tiers aren't fully separate states, this/that. Outside of the modest increases needed for compute, coding and dissemination, most funding above the minimum converts directly into protected research time. A partial commitment at any level between $12,000 and $62,000 produces proportional additional output.
For calibration/reference, v0.2 covered eight models across three vendors and produced a full public deposit with persistent DOIs, alongside responsible disclosure to Anthropic, OpenAI and Google. It was completed through out-of-pocket inference spending and unpaid evening work. What this funding does is alter/support how much of the work can be completed, maintained and released, and how quickly.