Apply the ADECP framework to audit and score frontier lab system cards, RSPs, and dangerous-capability disclosures, publishing a public report and ongoing tracker for comparable safety documentation.
Apply the ADECP framework to audit and score frontier lab system cards, RSPs, and dangerous-capability disclosures, publishing a public report and ongoing tracker for comparable safety documentation.
Project Details
Updated 07/08/26 · By grantmaking.ai · VerifiedI'm applying the AI Documentation Ethics Benchmark (ADECP), a framework I built and have used to grade consumer AI documentation for over a year, to a narrower and higher-stakes target: frontier lab system cards, responsible scaling policies, and dangerous capability evaluation disclosures. My MIT AI Safety Fellowship coursework covered deceptive alignment, mesa-optimization, and Anthropic's Sleeper Agents research, and one pattern kept surfacing: even when labs run genuinely rigorous internal safety evals, the public-facing documentation of those evals is often inaccessible to the people outside the lab who need to act on it, external safety researchers, journalists covering AI risk, and policymakers drafting oversight requirements. Funding would support a structured audit of system cards and safety disclosures from major labs (methodology adapted from my existing 7-Question Documentation Audit and four-category ADECP scoring), published as a public report and ongoing tracker. The theory of impact: legible, comparable disclosure of dangerous-capability findings is a precondition for the field noticing warning signs early and for regulators writing requirements that match actual practice rather than guesswork. Poor documentation is the bottleneck on the field's ability to respond to risk in time.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiReducing AI x-risk requires that warning signs get noticed and acted on before they compound into unrecoverable harm. That depends on safety-relevant information, dangerous capability eval results, red-teaming findings, deployment mitigations, actually reaching the people positioned to respond: external safety researchers who might catch something a lab missed, journalists who inform public and policymaker attention, and regulators writing binding requirements. Right now that information often exists but isn't legible. It's scattered across system cards, blog posts, and model cards using inconsistent structure, inconsistent depth, and inconsistent framing from lab to lab, which makes it hard to compare labs against each other or track whether disclosure is improving or degrading over time.
This project treats that legibility gap as a bottleneck. By applying a structured, repeatable audit methodology to frontier lab safety disclosures and publishing the results as a public, comparable tracker, the work does three things: it gives external researchers a faster way to spot which labs are under-disclosing specific categories of risk (rather than reading every system card cover to cover), it gives policymakers a evidence base for what "adequate disclosure" should require in binding regulation, grounded in actual practice rather than guesswork, and it creates public accountability pressure on labs to close specific, named gaps rather than vague ones.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.