AI risk assessment currently checks few threat models and doesn't compose them into the aggregate risk that matters. We'll build a tool mapping what frontier system cards cover and omit, plus a paper on what the assessments miss.
AI risk assessment currently checks few threat models and doesn't compose them into the aggregate risk that matters. We'll build a tool mapping what frontier system cards cover and omit, plus a paper on what the assessments miss.
Project Details
Updated 07/12/26 · Edited by orgCARMA will build a tool that maps the threat-model coverage of frontier AI system cards, run it on a handful of real system cards, and write up the results as a paper (for arXiv, and targeting a technical AI governance workshop, plus a companion blog post) on what those coverage gaps say about how AI risk is assessed today.
Frontier labs do look ahead in places. Their system cards report red-teaming and evaluations for a set of named harm categories, such as cyber and CBRN uplift, and with some a subset of RSI proxies. But that work covers a narrow part of the plausible threat surface and it stops at what the model can do (or is shown to be able to do) rather than what the model can be expected to do to the world it enters. A responsible assessment has to zoom out from the isolated model to its incremental effect on society: how its risks can be expected to propagate, how much it pushes the competitive race forward, how it weakens institutions, and how it shifts the defense burden society already carries. Almost no system card attempts this. Nor is any of it currently composed into a single view of the absolute risk that would actually justify, or fail to justify, the decision to train or deploy.
This is not a demand for the impossible. Reasoning about a technology's effects before those effects arrive is ordinary practice in mature safety-critical fields: no one proves a bridge will resonate to failure by building it and testing; they prove it beforehand, by theory, because you cannot afford to learn it by observation. While not fully comprehensive, much of the theory and methods needed to reason this way about AI and about societal risk propagation already exist, validated across subtopics like control theory, theory of agency, the physical and social sciences, and intelligent-systems dynamics. Our own published work shows a concrete route: Adapting Probabilistic Risk Assessment for AI (arXiv:2504.18536) is a top-down methodology that reaches a semi-quantified estimate of absolute risk, with judgment calls made explicit rather than hidden, and our offense-defense dynamics work (arXiv:2412.04029) orients such an assessment toward how a model shifts society's exposure. It is not a finished science, but it shows the zoom-out is doable. Though the relevant theory is somewhat fragmented across fields, it can and should inform how frontier systems are assessed. The obligation to carry that assessment out for a given model rests with the lab building it. This project checks whether they are meeting it.
The tool of the current proposal will make this concrete. Against a reference set of decision-relevant threat models, drawn from the MIT AI Risk Repository together with our published hazard taxonomy and our own more detailed threat models, and including the aggregate and systemic risks that per-model assessment tends to leave out, it maps what a given system card does and does not consider. For each threat model addressed, it records a coarse severity band (bounded, severe, or catastrophic/largely irreversible) and shows the reasoning behind that band rather than inventing a number. Coding will be LLM-assisted but human-verified. We'll apply it to a handful of current system cards and publish the resulting coverage maps, including in an interactive online format. The key findings will be the structured patterns of what is systematically left out, expected to particularly include aggregate, cross-actor, and long-tail catastrophic risks.
Things like SaferAI's ratings and the FLI AI Safety Index assess companies' published frameworks against governance criteria, and already document, in aggregate, that the reasoning linking evaluations to risk is usually thin and that existential safety planning is largely absent. Our tool will go a level down, to the per-release system cards and to the medium-fine-grained clusters of threat models, mapping which risks are actually addressed in the concrete disclosures on which deployment decisions rest.
Deliverables: 1) the coverage-mapping tool with 2) published results for several real system cards; 3) the paper, which sets out the families of evidence available for reasoning about risks that haven't fully materialized, shows how they can compose, and argues that assessment should both include the wider landscape of threat models and be organized around absolute risk (which is expected to highlight catastrophic risks more); and communication outputs, namely 4) a public explainer, 5) a one-page brief for practitioners and policymakers, and 6) a few explanatory figures. This is analytical work on public documents, divulges no sensitive threat model details, and is suitable for public release.
CARMA is a research organization working on the assessment and mitigation of catastrophic and existential risk from advanced AI. This project draws on our published research and the staff who produced it.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiWhether a frontier model gets trained at all, trained further, or deployed internally or externally increasingly turns on the lab's own risk assessment. When that assessment looks at a short list of harms, myopically and in isolation, the decision to keep scaling rests on evidence that can't speak to the society-wide, largely irreversible outcomes that make up most of the catastrophic risk. Those outcomes are not even necessarily properties of any single model in the first place but often live in the aggregate: many labs racing, dangerous capabilities spreading through open weights (which can be turbocharged by a closed-weight release), and capabilities recombining across systems and actors that no one evaluation watches. Assessing models one at a time is necessary, but a model's foreseeable marginal contribution to that aggregate, how far it pushes the race, how capabilities can recombine, what it does to institutions, how it shifts society's defensive burden, is left uncounted currently.
We want to move what practitioners, evaluators, and regulators treat as the thing being assessed from whether a model looks safe on today's tests, to what it adds to the absolute risk of societal harm and catastrophe once its effect on the wider system is counted. The tool we'll make will show this is doable. An argument about assessment methodology is easy to agree with and easy to shelve; a coverage map showing, for a specific system card, which risks and which societal effects went unexamined is something a lab, an evaluator, or a regulator has to either answer or ignore in public.
People
Updated 07/11/26 · Edited by orgTeam Member
Funding Details
- -
- -
- 5 months
- -
- -
- -
- -
- -
- seeking first grant for this project
- Social & Environmental Entrepreneurs
Track Record
CARMA (the Center for AI Risk Management and Alignment) is a research and policy thinktank working on the assessment, communication, mitigation, governance, and prevention of catastrophic and existential risks from advanced AI. We run relatively lean yet we engage dozens of institutions across government, academia, industry, and civil society in the US, Europe, and Asia.
Directly relevant published work:
"Adapting Probabilistic Risk Assessment for AI" (arXiv:2504.18536), a top-down methodology reaching a semi-quantified estimate of absolute risk for general-purpose AI. It has been referenced across sectors, from AI-safety organizations to academic governance groups to industry, and its first author now sits on the European Commission's AI Act Scientific Panel; the Commission's hosted biography describes her work as the "development of a probabilistic risk assessment methodology for general-purpose AI, which has informed the EU AI Act Code of Practice."
"Considerations Influencing Offense-Defense Dynamics from Artificial Intelligence" (arXiv:2412.04029), on factors affecting the balance between offense and defense societally across use cases and domains.
Standards and multi-institutional risk-management work, including co-authored Oxford Martin AI Governance Initiative memos on risk tiers and on open problems in frontier AI risk management.
Wider body of work:
We co-authored, with Yoshua Bengio, Audrey Tang, and Stuart Russell among others, a framework on how general-purpose AI threatens democratic and social systems. We published "Adaptive Governance for Advanced AI," introducing a diagnostic for distinguishing genuine oversight from governance that only performs it. We built and released public, running research software, including a multi-agent arbitration platform (supported by a Foresight Institute grant) and an LLM-assisted wargaming platform for crisis exercises. We published best practices regarding AI whistleblowing policy. We contributed significantly to the emergency-preparedness and prohibitions workstreams of a multilateral AI-governance treaty-design process, and have produced national-resilience and incident-response preparedness work for senior government audiences.
High impact work, bringing together existing work that is already established without reinventing the wheel. If implemented, it has the potential to solve coordination problem by joining hands with the key playes of the AI risk community.