Reducing AI x-risk requires that warning signs get noticed and acted on before they compound into unrecoverable harm. That depends on safety-relevant information, dangerous capability eval results, red-teaming findings, deployment mitigations, actually reaching the people positioned to respond: external safety researchers who might catch something a lab missed, journalists who inform public and policymaker attention, and regulators writing binding requirements. Right now that information often exists but isn't legible. It's scattered across system cards, blog posts, and model cards using inconsistent structure, inconsistent depth, and inconsistent framing from lab to lab, which makes it hard to compare labs against each other or track whether disclosure is improving or degrading over time.
This project treats that legibility gap as a bottleneck. By applying a structured, repeatable audit methodology to frontier lab safety disclosures and publishing the results as a public, comparable tracker, the work does three things: it gives external researchers a faster way to spot which labs are under-disclosing specific categories of risk (rather than reading every system card cover to cover), it gives policymakers a evidence base for what "adequate disclosure" should require in binding regulation, grounded in actual practice rather than guesswork, and it creates public accountability pressure on labs to close specific, named gaps rather than vague ones.