Project Details
Updated 07/10/26 · Edited by org- The problem: Frontier labs' current alignment strategy is primarily defense-in-depth - a stack of safety techniques at every layer from data collection to deployment. The number of safety techniques, published papers, AI models, expert opinions, etc. is vast and fast growing, and there is no place that collects this information and presents it in a comprehensive way - the closest existing resources are periodic prose overviews at the resolution of research agendas. Much of this information is unknown (e.g. performance in various contexts) or not public, and there is no "negative map" of what is unknown.
- The project: We will create a comprehensive catalogue of AI safety techniques (around 300 as of our prototype) to collect what is known and help expose where gaps may lie. The Stack will source and compile available public information including research papers, evidence of effectiveness and cost, implementations, deployment status across labs and models (often known only indirectly, if at all), blog posts and reports, expert opinions and critique, relationships and lineage; and aspirationally, mutual interactions of techniques.
- The output: A web-based catalogue of techniques, models, papers, labs, researchers, benchmarks, and expert opinions, with taxonomised and filterable views (e.g. timelines, epistemic coverage, estimates of promise and neglectedness) to see the state of the Stack. Every piece of information is sourced and carries epistemic flags (e.g. measured / self-reported / peer-reviewed / inferred / expert opinion / best guess). This enables a negative map as part of the output, recording what we looked for and did not find, and how thoroughly we looked.
- A substantial portion of the value lies in collecting and aggregating informal information and clearly marking them for their epistemic status: expert opinions, best guesses, notable anecdotal results, implied information, compliance filings, case proceedings, forecasts, etc. Informal sources like these are sometimes the only source of information about frontier models. The extracted, annotated data will be available as an open dataset (details pending licensing considerations).
- Primary users: safety researchers (orientation in the field), funders (neglected work, expert opinions, known usage; focus on philanthropic funding), and evaluators and policy analysts (as an overview of the stack components, and source index).
- Continuity: The collection will be updated ~monthly (adjustable). The pipeline is to be largely automated (AI-based extraction, classification and checks), with human oversight and spotcheck review, occasional code updates and maintenance. If this is successful, we expect to follow up with extension projects that will carry this in the medium-term, or with another grant.
- Prototype: We have explored this with an internal prototype: a taxonomy of ~300 techniques, ~2k sources, several experimental views over the extracted data. Based on this and our Shallow Review experience, we estimate ~30k source documents would cover most (>90%) of relevant public information.
- Project timeline: 3 months (1m to e2e pipeline; 1m full-data version, semi-public; 1m iterating final design with users, external reviews). Followed by 6-12 months of updates within this grant.
- Possible future extensions: Modelling technique effectiveness in different stacks, e.g. based on benchmarks. Extending the scope and structure of the data collected. An ambitious extension is performing - or contracting - experiments with the most informative stack combinations and gathering evidence on stack component interactions (with future R&D automation). Expert/researcher opinion surveys and larger-scale reviews of the data. Serving as a base for Shallow Review 2026.
- Example non-goals: Tracking individuals (beyond e.g. paper authorship). Judging labs' alignment efforts (like AI Safety Index). Calls to action other than research recommendation.
The team:
-
Tomáš Gavenčiak (project lead) led and built the 2025 Shallow Review of Technical AI Safety; a researcher at Alignment of Complex Systems research group, Charles University, Prague.
-
Jacob Livingston Slosser is building CheatSheet - a systematic catalogue of specification-gaming incidents in AI systems, and built Juriscription; a Law and AI scholar at the Pioneer Centre for Artificial Intelligence (P1) and the University of Copenhagen.
-
Dan Elton built the Metascience Observatory, a living annotated map of the metascience literature.
-
Gavin Leech (advisor) built the earlier iterations of the Shallow Review, along with other projects; director of Arb Research. (Gavin has a declared CoI with grantmaking.ai; here in an unpaid advisory role.)
Theory of Impact
Updated 07/10/26 · By grantmaking.aiThe deployed stack of safety techniques is a large part of how AI risk is actually being managed. Making it visible improves safety efforts along the following routes:
- Better allocation: Current allocation of research, funding, and attention is only partially based on deep expertise, and to a large extent follows trends and salience. We will provide both comprehensive views (including negative spaces and neglected areas), and concrete, sourced evidence for its claims. This includes improved orientation in the field (researchers: existing and new research; non-specialists: existing techniques and their properties). Note that this project does not aim to replace expert judgment on area prioritization but rather complement it and record expert opinions.
- Epistemic pressure: Legibly demonstrating the extent of unknowns, non-public information and self-reporting - as well as creating spaces that track and report the epistemic status - creates incentives: for funding specific research and replications, for wider stack disclosure from the labs, and towards policy attention to the gaps. This seems valuable even at modest effect sizes.
People
Updated 07/10/26 · Edited by orgTeam Member
Track Record
See the team above!
Discussion
My intuition is that this stack would help a neglected area by serving those existing folks in government, politics, policy facing into wrapping their head around the literature, as they would be at a disadvantage in terms of research taste and judgment than those in the field. regardless of that use case, this seems worth supporting.
Hi Tomáš,
We'd like to fund this for $50k (our maximum). Logistical questions:
-
Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
-
Please confirm your commitment to post quarterly updates on how the project is going
- How will funding be disbursed? Equal split to three individuals?
I'm very excited about this. The opacity of the "stack" to roughly everyone in the world and many staff within frontier labs strikes me as a full-blown epistemic crisis and a frankly illegitimate state of affairs. So besides the technical benefits (paying down a vast pile of research debt and drawing attention to the actually important context of evaluation, viz. the joint effect of safety techniques), I view this as an important intervention into technical policy. People need to know how unknown this is, and people need to know the thinness of the stack without letting obscurity comfort them by implying that secret alignment tech will catch us as we fall.
A major question I have is whether you can go beyond collating public information, towards clever inference of techniques that must be being used, OSINT, or gaining the trust of whistleblowers.
One last major question I have is: what are the dual-use implications?
Good luck!
Conflict of interest: Tomas is a close friend and I helped conceive the project. Obviously I won't take any funding myself. I also wouldn't fund this if it hadn't gotten a strong endorsement from another grantmaker.
That is great news - thank you!
1. Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
No to both.
2. Please confirm your commitment to post quarterly updates on how the project is going
We do. Say 16 Sept latest, though I hope we'll have a shareable prototype much sooner.
3. How will funding be disbursed? Equal split to three individuals?
Would prefer invoicing and flexibility to contract others, so looking for a EA-adjacent fiscal sponsor (invoices, reimbursement, accounting). Discussing details with Matt.
The first phase is mapping relevant literature (+LW/AF, blog posts, ...) to alignment techniques (a rich taxonomy of techniques is WIP), models mentioned/used/..., and orgs&labs. The first goal is presenting this as a technique/literature survey for both alignment researchers and experts in policy, govt and journalism. I agree that this already seems useful @Pip Foweraker! As a reference point, the Shallow Review of Technical AI safety is our earlier project in that direction, though primarily intended for researchers.
The overall goal is to include information about the deployment of alignment techniques, though this information is rarely known for closed models beyond model&system cards, and researcher&expert opinions and evaluations of the techniques. We are looking into indirect sources of information, e.g. lab/personnel/technique/open-implementation incidence, expert opinions and predictions, insider comments, and some less common public sources, e.g. legal and regulatory filings - we have to look into them and estimate the useful information.
To answer @Gavin Leech, both OSINT and whistleblower information seem great, but may not fit this phase.
Everyone: if you are have insight into OSINT information sources relevant to this area, we'd appreciate it if you reached out to us!
Dual-use: Likely low since this is primarily public-data based, and the labs have their own experts and private know-how, but taking this seriously: This may help orientation in dual-use alignment techniques (i.e. where performance improves with alignment). In particular, collecting benchmark scores across models and techniques could be used to infer capability gains (as well as costs) of techniques at scale. Unlikely to help top labs but could potentially help the wider ML community (academia, startups). Mitigation is focus on alignment techniques.
Private comment. Only shown to approved funders and grant reviewers.