grantmaking.ai Launch Round
The problem: Frontier labs' current alignment strategy is primarily defense-in-depth - a stack of safety techniques at every layer from data collection to deployment. The number of safety techniques, published papers, AI models, expert opinions, etc. is vast and fast growing, and there is no place that collects this information and presents it in a comprehensive way - the closest existing resources are periodic prose overviews at the resolution of research agendas. Much of this information is unknown (e.g. performance in various contexts) or not public, and there is no "negative map" of what is unknown.
The project: We will create a comprehensive catalogue of AI safety techniques (around 300 as of our prototype) to collect what is known and help expose where gaps may lie. The Stack will source and compile available public information including research papers, evidence of effectiveness and cost, implementations, deployment status across labs and models (often known only indirectly, if at all), blog posts and reports, expert opinions and critique, relationships and lineage; and aspirationally, mutual interactions of techniques.
The output: A web-based catalogue of techniques, models, papers, labs, researchers, benchmarks, and expert opinions, with taxonomised and filterable views (e.g. timelines, epistemic coverage, estimates of promise and neglectedness) to see the state of the Stack. Every piece of information is sourced and carries epistemic flags (e.g. measured / self-reported / peer-reviewed / inferred / expert opinion / best guess). This enables a negative map as part of the output, recording what we looked for and did not find, and how thoroughly we looked.
A substantial portion of the value lies in collecting and aggregating informal information and clearly marking them for their epistemic status: expert opinions, best guesses, notable anecdotal results, implied information, compliance filings, case proceedings, forecasts, etc. Informal sources like these are sometimes the only source of information about frontier models. The extracted, annotated data will be available as an open dataset (details pending licensing considerations).
Primary users: safety researchers (orientation in the field), funders (neglected work, expert opinions, known usage; focus on philanthropic funding), and evaluators and policy analysts (as an overview of the stack components, and source index).
Continuity: The collection will be updated ~monthly (adjustable). The pipeline is to be largely automated (AI-based extraction, classification and checks), with human oversight and spotcheck review, occasional code updates and maintenance. If this is successful, we expect to follow up with extension projects that will carry this in the medium-term, or with another grant.
Prototype: We have explored this with an internal prototype: a taxonomy of ~300 techniques, ~2k sources, several experimental views over the extracted data. Based on this and our Shallow Review experience, we estimate ~30k source documents would cover most (>90%) of relevant public information.
Project timeline: 3 months (1m to e2e pipeline; 1m full-data version, semi-public; 1m iterating final design with users, external reviews). Followed by 6-12 months of updates within this grant.
Some non-goals: Tracking individuals (beyond e.g. paper authorship). Judging labs' alignment efforts (like AI Safety Index). Calls to action other than research recommendation.
The team:
-
Tomáš Gavenčiak (project lead) led and built the 2025 Shallow Review of Technical AI Safety; a researcher at Alignment of Complex Systems research group, Charles University, Prague.
-
Jacob Livingston Slosser is building CheatSheet - a systematic catalogue of specification-gaming incidents in AI systems, and built Juriscription; a Law and AI scholar at the Pioneer Centre for Artificial Intelligence (P1) and the University of Copenhagen.
-
Dan Elton built the Metascience Observatory, a living annotated map of the metascience literature.
-
Gavin Leech (advisor) built the earlier iterations of the Shallow Review, along with other projects; director of Arb Research. (Gavin has a declared CoI with grantmaking.ai; here in an unpaid advisory role.)
Possible future extensions: Modelling technique effectiveness in different stacks, e.g. based on benchmarks. Extending the scope and structure of the data collected. An ambitious extension is performing - or contracting - experiments with the most informative stack combinations and gathering evidence on stack component interactions (with future R&D automation). Expert/researcher opinion surveys and larger-scale reviews of the data. Serving as a base for Shallow Review 2026.
The deployed stack of safety techniques is a large part of how AI risk is actually being managed. Making it visible improves safety efforts along the following routes:
Better allocation: Current allocation of research, funding, and attention is only partially based on deep expertise, and to a large extent follows trends and salience. We will provide both comprehensive views (including negative spaces and neglected areas), and concrete, sourced evidence for its claims. This includes improved orientation in the field (researchers: existing and new research; non-specialists: existing techniques and their properties). Note that this project does not aim to replace expert judgment on area prioritization but rather complement it and record expert opinions.
Epistemic pressure: Legibly demonstrating the extent of unknowns, non-public information and self-reporting - as well as creating spaces that track and report the epistemic status - creates incentives: for funding specific research and replications, for wider stack disclosure from the labs, and towards policy attention to the gaps. This seems valuable even at modest effect sizes.
Groundwork for understanding technique interactions and real-world performance: Defense-in-depth assumes that its layers fail mostly independently; this is hard to study without the stack being extensively mapped. Understanding the stack as a whole in terms of technique interactions and risks countered by its layers would enable better resource allocation, deployment decisions, and regulation. Aspirational and exploratory, but high potential value even at only a small contribution; a substantial fraction of our original motivation.
Counterfactual projects: Literature search tools - great for concrete queries but no aggregate lens. Model&system cards (only specific information). Shallow Review of Technical AI Safety (our own project) - different focus (agendas; for researchers), fewer details, 1/y project as of now. FLI's AI Safety Index and AISI's Frontier AI Trends Report (focus on big picture and high-level policy implications).
Risks: Biased aggregations and high-level views presented with overconfidence, or misinterpreted by researchers, funders, or journalists - mitigated by epistemic status, explicit flagging, and external reviews (e.g. sanity checks on the high-level implied messaging). Failed fit with target users - mitigated by iteration, and by our experience with Shallow Review and other projects (fallback to decreased project value rather than failure).
About 80% of the budget is core team salaries: 3 months of 1 FTE (average over the team) to design, implement, run, and review the project, including 9 months of updating the dataset and code maintenance. Other expenses include compute (LLM API costs, $5k), external reviews (about 50h), <5% admin overhead. (Detailed budget shared privately.)
Private comment. Only shown to approved funders and grant reviewers.