grantmaking.ai Launch Round
AI risk assessment currently checks few threat models and doesn't compose them into the aggregate risk that matters. We build a tool mapping what frontier system cards cover and omit, plus a paper on what the assessments miss.
CARMA will build a tool that maps the threat-model coverage of frontier AI system cards, run it on a handful of real system cards, and write up the results as a paper (for arXiv, and targeting a technical AI governance workshop, plus a companion blog post) on what those coverage gaps say about how AI risk is assessed today.
Frontier labs do look ahead in places. Their system cards report red-teaming and evaluations for a set of named harm categories, such as cyber and CBRN uplift. But that work covers a narrow part of the plausible threat surface and it stops at what the model can do (or is shown to be able to do) rather than what the model can be expected to do to the world it enters. A responsible assessment has to zoom out from the isolated model to its incremental effect on society: how its risks can be expected to propagate, how much it pushes the competitive race forward, how it weakens institutions, and how it shifts the defense burden society already carries. Almost no system card attempts this. Nor is any of it currently composed into a single view of the absolute risk that would actually justify, or fail to justify, the decision to train or deploy.
This is not a demand for the impossible. Reasoning about a technology's effects before those effects arrive is ordinary practice in mature safety-critical fields: no one proves a bridge will resonate to failure by building it and testing; they prove it beforehand, by theory, because you cannot afford to learn it by observation. While not fully comprehensive, much of the theory and methods needed to reason this way about AI and about societal risk propagation already exist, validated across subtopics like control theory, theory of agency, the physical and social sciences, and intelligent-systems dynamics. Our own published work shows a concrete route: Adapting Probabilistic Risk Assessment for AI (arXiv:2504.18536) is a top-down methodology that reaches a semi-quantified estimate of absolute risk, with judgment calls made explicit rather than hidden, and our offense-defense dynamics work (arXiv:2412.04029) orients such an assessment toward how a model shifts society's exposure. It is not a finished science, but it shows the zoom-out is doable. Though the relevant theory is somewhat fragmented across fields, it can and should inform how frontier systems are assessed. The obligation to carry that assessment out for a given model rests with the lab building it. This project checks whether they are meeting it.
The tool of the current proposal will make this concrete. Against a reference set of decision-relevant threat models, drawn from the MIT AI Risk Repository together with our published hazard taxonomy and our own more detailed threat models, and including the aggregate and systemic risks that per-model assessment tends to leave out, it maps what a given system card does and does not consider. For each threat model addressed, it records a coarse severity band (bounded, severe, or catastrophic/largely irreversible) and shows the reasoning behind that band rather than inventing a number. Coding will be LLM-assisted but human-verified. We'll apply it to a handful of current system cards and publish the resulting coverage maps, including in an interactive online format. The key findings will be the structured patterns of what is systematically left out, expected to particularly include aggregate, cross-actor, and long-tail catastrophic risks.
Things like SaferAI's ratings and the FLI AI Safety Index assess companies' published frameworks against governance criteria, and already document, in aggregate, that the reasoning linking evaluations to risk is usually thin and that existential safety planning is largely absent. Our tool will go a level down, to the per-release system cards and to the medium-fine-grained clusters of threat models, mapping which risks are actually addressed in the concrete disclosures on which deployment decisions rest.
Deliverables: 1) the coverage-mapping tool with 2) published results for several real system cards; 3) the paper, which sets out the families of evidence available for reasoning about risks that haven't fully materialized, shows how they can compose, and argues that assessment should both include the wider landscape of threat models and be organized around absolute risk (which is expected to highlight catastrophic risks more); and communication outputs, namely 4) a public explainer, 5) a one-page brief for practitioners and policymakers, and 6) a few explanatory figures. This is analytical work on public documents, divulges no sensitive threat model details, and is suitable for public release.
CARMA is a research organization working on the assessment and mitigation of catastrophic and existential risk from advanced AI. This project draws on our published research and the staff who produced it.
Whether a frontier model gets trained at all, trained further, or deployed internally or externally increasingly turns on the lab's own risk assessment. When that assessment looks at a short list of harms, myopically and in isolation, the decision to keep scaling rests on evidence that can't speak to the society-wide, largely irreversible outcomes that make up most of the catastrophic risk. Those outcomes are not even necessarily properties of any single model in the first place but often live in the aggregate: many labs racing, dangerous capabilities spreading through open weights (which can be turbocharged by a closed-weight release), and capabilities recombining across systems and actors that no one evaluation watches. Assessing models one at a time is necessary, but a model's foreseeable marginal contribution to that aggregate, how far it pushes the race, how capabilities can recombine, what it does to institutions, how it shifts society's defensive burden, is left uncounted currently.
We want to move what practitioners, evaluators, and regulators treat as the thing being assessed from whether a model looks safe on today's tests, to what it adds to the absolute risk of catastrophe once its effect on the wider system is counted. The tool we'll make will show this is doable. An argument about assessment methodology is easy to agree with and easy to shelve; a coverage map showing, for a specific system card, which risks and which societal effects went unexamined is something a lab, an evaluator, or a regulator has to either answer or ignore in public.
The standing objection would be that putting a number on catastrophes with no track record just invents numbers, and that testing the current model is the only rigorous move. That has the direction of rigor backwards. Counting only what you can measure on the system in front of you is not caution but rather a category error about the question, because the question is about effects that by definition mostly haven't happened yet. Every mature safety-critical field reasons about failures it cannot afford to observe, using theory, analogy, and disciplined judgment alongside measurement, with the assumptions written down and open to attack and iteration; too much is known of the risk dynamics to just dismiss it all as pre-paradigmatic. Our own probabilistic risk assessment work shows this is tractable for AI in a best-effort manner: a top-down estimate that is semi-quantified and explicit about where judgment enters, rather than either false precision or hand-waving. The tool in the current scope holds itself to the same standard, reporting severity as reasoned bands with their basis shown. It is not asking labs to invent numbers; it is asking them to reason where they currently go silent.
What counts as an adequate assessment propagates into safety cases, evaluation standards, and the regulation built on them. Fixing the target of assessments is therefore worth more than fixing any one assessment. Since such a reframe only changes practice if the people who set expectations actually run into it, the paper will travel with a plain-language explainer, a one-page brief, and figures appropriate for the salient audiences.
Minimum: $38,000 (about 2.4 FTE-months over roughly 5 months). This covers the core deliverables: the reference threat-model set, the coverage-mapping tool applied to a small handful of real system cards with published results, the paper, a public explainer, and a one-page brief with explanatory figures.
The bulk is senior researcher time, about 0.45 FTE over 5 months (2.25 FTE-months), for assembling the reference threat-model set, human-verified coding of the system cards, developing the argument, and writing the paper. That's $33,750 direct. Research assistance and production (tool scaffolding, the LLM-assisted coding pipeline, coverage-map figures, blog adaptation) adds 0.15 FTE-months, or about $1,900. Direct costs come to roughly $35,650; with 6.5% overhead (roughly $2,350), the total is approximately $38,000.
Ideal: $62,000 (about 3.9 FTE-months over roughly 8 months). Everything above, applied to more system cards and with a fuller public coverage-map visualization, plus a worked application of the lens to several published safety frameworks, additional audience-tailored briefs, outside expert review, and workshop submission.
The additional funding covers: 1.3 more FTE-months of research time for the extra cards, more granular threat models, fuller visualization, framework application, expanded communication materials, and revision ($19,500 direct); honoraria for two to three external reviewers from the risk-assessment and evaluation community (roughly $2,500); and workshop submission, registration, and modest travel contingent on acceptance (roughly $1,700). Direct costs come to approximately $58,200; with overhead (roughly $3,800), the total is approximately $62,000.
Note that we budgeted researcher time at a loaded rate of $180k per FTE-year ($15k per FTE-month).
Between the minimum and ideal amounts, the additional funding extends the tool to more system cards, adds granularity to the threat models addressed, broadens the communication materials and their reach, and widens expert review.
High impact work, bringing together existing work that is already established without reinventing the wheel. If implemented, it has the potential to solve coordination problem by joining hands with the key playes of the AI risk community.