grantmaking.ai Launch Round
Most concrete AI loss-of-control stories pass through a moment where a human failed at oversight: someone let an autonomous agent keep running after it exceeded its brief, trusted a fluent but false account of what a model did, accepted a reassuring answer under pressure, or missed that a system was acting with authority nobody granted it.
Many deployment strategies for increasingly autonomous AI systems rely on exactly this kind of human supervision as a final safety backstop. They assume a person will notice when an agent behaves unsafely, establish what actually happened and intervene before the consequences become difficult to reverse. As AI systems get more agentic, this human oversight layer becomes load-bearing and it is badly under-studied. The field invests heavily in evaluating models; almost nothing measures where humans actually break down when supervising them, or trains that skill on purpose.
This safety assumption can fail at several distinct points and each one may call for a different safeguard. Oversight Lab evaluates and distinguishes three:
- Detection: relevant warning signs are missed entirely.
- Verification: a supervisor notices something unusual but doesn't test the agent's explanation against independent evidence.
- Intervention: a supervisor correctly identifies unsafe behaviour but still fails to restrict, suspend or escalate it in time.
We already have a working, playtested prototype that puts this taxonomy to the test. Careful supervision, verifying evidence, scoping authority, escalating in time, scores near-perfectly across every oversight competency we grade for; careless supervision fails every one of them. That gap is the first evidence that the evaluation actually discriminates safe from unsafe behaviour, rather than rewarding luck or familiarity with the interface.
This distinction matters because it points to different fixes. Detection failures suggest a need for clearer alerts or more legible activity records. Verification failures suggest a need for independent logs, provenance checks, or easier access to contradictory evidence. Intervention failures suggest a need for predefined escalation thresholds, clearer responsibilities, or more reliable ways to restrict and suspend an agent. Knowing which failure mode is actually driving a supervision breakdown is what lets a safety team fix the right thing instead of guessing.
The causal pathway is: behavioural evaluation → evidence on whether oversight fails at detection, verification or intervention → evidence-based improvements to control mechanisms and deployment processes → more reliable human oversight → a reduced likelihood that unsafe agentic behaviour continues unchecked.
The connection to existential risk is indirect but concrete. A failure in one simulation is not itself an existential event. The concern is that the same classes of human failure, missed signals, unverified claims, delayed intervention, may persist as AI systems become more capable, more persuasive, more interconnected and able to take increasingly consequential actions on their own. If supervisors cannot reliably detect unsafe behaviour, establish what an agent actually did, or intervene in time, deployment strategies that depend on human oversight provide substantially less protection than they assume.
Oversight Lab aims to surface these weaknesses now, while human supervision is still a plausible line of defence, before it is relied upon for systems that are substantially harder to control. Its open framework and dataset will also let other researchers examine which forms of human oversight are actually reliable and where additional technical or organisational safeguards are needed.
Oversight Lab is not itself a scalable-oversight protocol. Its direct contribution is human-side evaluation and AI control; its findings may inform future work on scalable oversight and AI-assisted oversight.
Simulation results will not perfectly predict behaviour in every real-world deployment. External review, broader testing and transparent reporting of methodological limitations are therefore part of the proposed work, not an afterthought.
IDEAL ($50,000) for:
-
Research & expert review (~$12k): paid consultations with AI-safety
researchers to ground new scenarios in real agentic-failure modes
(specification gaming, deceptive/unfaithful reasoning, sycophancy, loss
of control); literature synthesis; adversarial review of scenarios for
fidelity, so they teach the real failure mode and not a caricature of it. -
Content generation & scenario development (~$15k): expanding the
scenario library from 1 to 8-12 agentic-AI-safety scenarios, bilingual,
with realistic in-story artifacts; AI-assisted authoring with human
editing and adversarial playtesting of every scenario before
release; designing and validating the human-oversight eval instrument and its scoring. -
Development & productionisation (~$12k): turning the current prototype
into a reliable and accessible public instance, the anonymised free
eval-export pipeline, dataset publication tooling, accessibility work,
and self-serve access for researchers/educators/individuals. -
Compute & hosting (~$5k): LLM API costs for live simulations, AI-assisted
content generation, and eval runs; hosting the free public instance for
12 months. -
Dissemination & outreach (~$6k): getting the training and the free open
dataset in front of the safety research community, educators, and
practitioners who supervise AI systems day to day; a small public page
for the dataset; the quarterly public updates.
MINIMUM ($20,000) for:
The version that still ships something real if we're underfunded: a first round of research-reviewed agentic-safety scenarios (2-4 new, alongside the existing one), the public instance and dataset v1 published, cut from content depth and outreach, not from the public-benefit deliverables themselves.
Both tiers include the founder's own (AI-assisted) labor; AI leverage in research synthesis, scenario drafting, and playtesting is why this budget funds a program, not a single feature.