This project evaluates model cards and related benchmarks to determine the quality of reporting and estimating the life span of benchmarks before saturation.
This project evaluates model cards and related benchmarks to determine the quality of reporting and estimating the life span of benchmarks before saturation.
Project Details
Updated 07/22/26 · Provided via application · VerifiedOrganizations & Objectives:
The University of Oslo (UiO) and Frankfurt AI Safety are involved in this research.
The objective of this work is to evaluate model cards from frontier organizations and determine the quality and lifespan of benchmarks to determine the current state of red teaming and safety.
Output:
The output will be a research paper. We are targeting, "19 th ACM Workshop onArtificial Intelligence and Security." This co-located with the 33rd ACM Conference on Computer and Communications Security.
Team:
Ryan Marinelli - leading red teaming for this project. He is using functional data analysis to estimate the lifespan of benchmarks before they become saturated.
Robert Andrew Chetwyn - supporting argumentation with his cybersecurity background to organize using reports from the cybersecurity community to triangulate benchmarks.
Cedric Kopp - gathering data from model cards and related benchmarks to support the analysis. He is also supporting efforts in data visualization and other analysis.
Sara Hobe - applies qualitative frameworks to benchmarks to determine the validity of the benchmarks studied.
Helen Adepoju - supporting in project management and scheduling between different parties.
Theory of Impact
Updated 07/22/26 · By grantmaking.aiOrganizations increasingly rely on model cards to make real decisions: procurement, deployment, risk assessments. Those decisions are only as good as the benchmark numbers underneath them. When a model card cites a saturated or invalid benchmark without flagging it, decision-makers are misinformed in a way that compounds. One bad call based on an inflated score doesn't stay isolated; it propagates into downstream systems and policies built on top of that assumption, creating systematic failure risk rather than a single point failure.
Triangulation and validation address this by making benchmark reliability legible to the reader rather than assumed. Instead of a model card presenting a score as face-value truth, this work lets a reader weigh how much confidence that score actually deserves, turning benchmark reporting from opaque to transparent and giving decision-makers the discernment to catch a stale or gamed number before they act on it.
The project also has a forward-looking arm: producing concrete recommendations for how benchmarks should be designed to resist saturation in the first place, rather than only diagnosing decay after the fact.
One detail worth calling out on the cybersecurity angle specifically: using community threat-reporting isn't just about better ground-truth for cyber-capability scores. It's a deliberate buffer against evaluation awareness. A model that can detect it's being tested can behave differently during a benchmark than it would in a real deployment, which would corrupt saturation estimates based on benchmark performance alone. Real-world cybersecurity community reports come from actual incidents and behavior in the wild, not from the model performing for an eval, so they're far harder for a model to game even if it recognizes it's being benchmarked elsewhere.
People
Updated 07/22/26 · By grantmaking.aiTeam Member
Funding Asks
Discussion
No comments yet. Be the first to share your thoughts.