*Comprehension audits* are a novel development-process assurance mechanism to verify human understanding of AI research outputs to act as a gate to slow automation of AI R&D.
*Comprehension audits* are a novel development-process assurance mechanism to verify human understanding of AI research outputs to act as a gate to slow automation of AI R&D.
Project Details
Updated 07/13/26 · Provided via application · VerifiedWe propose an auditing system to supply that missing gate. Comprehension audits are a novel development-process assurance mechanism to verify human understanding of AI research outputs. They test whether the responsible human teams can explain specific aspects of frontier AI R&D. With independent administration and graded reports, they provide a gate to stop based on a failure to demonstrate human understanding until remediated. They are complementary to most existing safety mitigations that examine research artifacts. They are also distinct from scalable oversight that targets the safety risk from AI outputs too advanced for human understanding.
Comprehension audits have evaluators interrogate human experts to verify their understanding of a key research output. This acts as a brake on recursive self-improvement: automated AI research can proceed only as fast as humans can demonstrably understand it, so a failure forces an organization to slow down and re-establish understanding before continuing. We advocate for labs adopting internal audits but incorporating external audits as they adopt AAL-2 or higher Assurance Levels (as defined by Brundage et al.).
The project also proposes monitoring automated metrics as indications that human oversight is declining, including the amount of time spent in reviewing pull requests by human reviewers, changes in the number of human reviewers, and the quality of human feedback. The project defines a methodology and standards for performing comprehension audits, situates it within the existing audit regime, and analyzes the impact on risks for automated AI R&D.
Our goal is to pilot the technique to demonstrate efficacy and to refine the methodology.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiAI is already writing a majority of code for frontier AI labs. This creates a safety risk from insufficient oversight due to time pressure. Existing work proposes minimum comprehension thresholds and unaided checks. To our knowledge, however, there is currently no published frontier-AI assurance regime that requires demonstrated evidence that the responsible humans understand what they are building a precommitted condition for continuing development or deployment.
RSI is the biggest accelerant of x-risk and establishing viable mechanisms to slow it down is a critical priority.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
Hi Gavin - here's a description with more details from the draft paper. I don't yet have a worked example but I could follow up later this week with one if that would still be helpful beyond the description (we will plan to add an example to the paper as an appendix):
Audit Meeting: The audit meeting is similar in nature to an oral examination in a thesis defense or to a design review. The auditors' agenda is used to pose questions to contributors. They are asked to explain detailed aspects of the algorithms and code and results produced based on them. The contributors are allowed to use tools to navigate and display work but not to do analysis or generate responses. They are expected to be able to synchronously answer questions. The effort required to answer is an important consideration for the auditor as is accuracy and thoroughness of response.
If a topic comes up in the audit review where there is not a human expert, the auditors can request that an appropriate expert joins the review meeting or that a follow-up meeting with one is scheduled. The auditors may also curtail a line of inquiry that isn't covered by the experts of the contribution. In general, comprehension audits should be topically focused on the work and supporting knowledge of those doing that work so this should be a relatively rare occurrence. If the responsible contributor for a part of the contribution has left the organization or due to recusal the auditor should simply verify that the departed contributor was indeed responsible from historical evidence. However, if the contribution actively extends work of a departed contributor, those making continued contributions based on the work must demonstrate knowledge within a precommitted period based on the criticality of the contribution.
Meetings should be scheduled for a defined maximum duration (e.g., ninety minutes). Auditors should have latitude to deviate from their agenda to follow up on answers, including new topics based on those answers (e.g., where an answer suggests a lack of understanding). This is also important to ensure that the audit can probe for rehearsed answers or other indications of cramming instead of true understanding.
The meetings should be recorded to support accurate reporting and any appeals. This minimally will require contemporaneous minutes, although electronic recording with automated transcription is the norm.
The framing of the audit meeting is intended to be that of a blameless postmortem: it is to test for adequate understanding and to allow for remediation of any lack of understanding. It is highly desirable that failures of understanding do not result in disciplinary action or negatively affect performance reviews by participants.
Grading:
The auditors will evaluate the level of understanding against rubrics for depth of understanding and identify areas of deficiency. These will typically be scoped for the contribution under review as a whole but may be broken into components or areas such as subsystem or domain (e.g., pretraining, post-training, and control system). Understanding will be assessed for topical areas:
Topical areas assessed in grading:
Architecture and design: The overall software components, how they work together, what key decisions were made in terms of how they would work and cooperate, optimization approach and what datasets were used. Includes factual recall, abiilty to modify or intervene, ability to explain motivations. Includes awareness of boundaries and dependencies outside core contribution.
Software implementation: What choices were made in implementing key parts of the contribution, including algorithms, security, scalability. The auditors are expected to identify areas of code of interest based on complexity, novelty or evidence of poor review to investigate. Includes factual recall, abiilty to modify or intervene, ability to explain motivations
Process: What process was followed, why were changes made, what were key tradeoffs considered in design and during experiments that motivated changes or updates. Includes factual recall, abiilty to modify or intervene, ability to explain motivations, predict consequences of process changes. The auditors are expected to identify important changes and updates made to probe during the audit.
Key findings: The overall results of experiments or implementations and how they were derived. Includes factual recall, explain causes, predict consequences, boundary awareness and calibrated uncertainty.
Analysis of results: Results of evaluations performed, any metrics computed, any failures encountered and reasons for them, and any anomalies identified and updates in response to them. Includes factual recall, explain causes, predict consequences, boundary awareness and calibrated uncertainty.
The expectation is that the software code, architecture, training setup, datasets used, and experiment goals will be deeply understood by the human contributors. They should be familiar with metrics and evaluations of the contribution and able to provide evidence of having investigated them to identify any discrepancies aligned with best practices for AI research. The human contributors are not expected to understand model internals, to explain why models acted in a different way, or otherwise exceed the level of understanding of deep learning that prevails in the industry. The audit assesses collective knowledge by the responsible team, not that of individual members. For any important aspect of the contribution at least one person must be able to adequately respond. Different contributors may respond to different questions and aspects and no advanced designation is required. The audit should weight evidence that understanding existed when the work was approved, not understanding that was remediated in preparation for or during an audit.
The auditors should consider anxiety, disabilities and language fluency (especially for non-native speakers) as mitigating factors for answers and allow respondents time and space to effectively answer. The blameless culture as noted earlier in the section should mitigate some anxiety. The purpose of the audit is not to assess fluency or confidence but understanding.
The grade will be either a pass, a conditional pass with remediation conditions or a failure. A single critical knowledge gap or a pattern of inadequate understanding in any important subdomain is sufficient to warrant a failure. Otherwise, the grading will weight the level of understanding against the auditor assessment of importance of the artifact or domain to produce an overall assessment of understanding from the audit.
Audit grades and their criteria:
Pass: Demonstrating minimal acceptable understanding in all components and aspects examined. There may be acceptable but suboptimal knowledge gaps or other areas that could be improved that are addressed in optional recommendations. Acceptable understanding should be found for a component if both of the following are true: (1) for any questions about all deep understanding tier aspects (architecture and design, software implementation, and process) at least one contributor is able to adequately answer them; and (2) for both familiarity tier aspects (key findings and analysis of results) for any question, at least one contributor is able to demonstrate due consideration.
Conditional pass: A failure in a noncritical component or area may be assigned a conditional pass with a requirement to improve understanding and/or adoption of changed processes to improve understanding. If there is a lack of knowledge due to staff turnover in an area during a transition period, a conditional pass would require a current team member to obtain sufficient knowledge in the area.
Failure: Any significant lack of understanding of important concepts or a general pattern of insufficient understanding in an important component or central aspect of a contribution will lead to a failure.
Audit outputs:
The auditors produce a written report after the audit meeting, with an overall grade plus any required remediation. The report will include a breakdown of grading of understanding by topical area. For larger contributions, this can be further decomposed into subdomains. The report will include a series of claims about the level of understanding in each topical area with reference to evidence from the audit meeting as well as supporting evidence from analysis of the R&D artifacts and metrics.
The audit report will note knowledge gaps. Where there is evidence that an expert contributor was not in attendance, a lack of comprehension by those in attendance is weak evidence of a lack of understanding by current staff. However, a frequent appeal to others who are no longer at the organization or are in a supporting role but not invited to the audit is itself evidence of a lack of comprehension. The auditors will need to judge based on the knowledge exhibited by those present, the pattern of contribution from the historical data, and the time that has elapsed since a contributor left the organization.
The report will specify any required remediation for insufficient understanding as well as including any optional recommendations for improvement.
These scenarios were generated by Claude Fable with editing by the principal author. Both are fictional composites constructed to illustrate the examination method and the grading rubrics in Appendix A. They do not describe any actual organization, team, or audit.
The two cases below show examples of audits where there are meaningful findings, to test the boundary of the rubric. The first is a comprehension failure that surfaces only under a follow-up question, after architecture-level answers had passed. The second shows conduct that might trouble an auditor on first impression, a shortcut taken under deadline pressure, that grades as a pass because the team demonstrates understanding of the system, including the part it chose to defer. These examples highlight how audits are intended to discriminate human understanding, not as an assessment of overall research quality or best practices. Routine opportunities for performance improvement can result in advisory recommendations but are not deemed to be comprehension failures.
Case A: an implementation failure behind passing architecture answers
Setup. The audited artifact is an automated experiment harness developed by a five-person applied research team. The harness sweeps fine-tuning configurations, runs each configuration in an isolated sandbox, scores completed runs with a composite metric, and automatically promotes the best-scoring checkpoint to a shared evaluation queue used by other teams. Roughly 70 percent of the harness code was generated by coding agents over six weeks. Process metrics for the contribution were healthy throughout: CI green, every PR reviewed, review latency normal. The audit team selected the contribution after monitoring metrics showed the team's merged volume tripling while review comments per line fell. The team lead was notified of the audit and he also required that the two largest contributors attended the meeting. Preparation was unannounced and the examination followed the no-AI rule.
Excerpted exchange (condensed).
Auditor: Walk me through the harness end to end.
Lead: The sweep controller reads the configuration grid, spawns one sandbox per configuration, and streams metrics to the tracker. When the sweep completes, the scorer computes a composite of eval accuracy, regression-suite pass rate, and a safety-filter score, and the top checkpoint is promoted to the shared queue.
Auditor: Why a composite score rather than gating on each metric separately?
Lead: We wanted a single ranking so promotion could be automatic. Separate gates stalled too many sweeps. The composite weights were tuned against six historical sweeps.
Auditor: How does the composite handle a run that terminates early?
Lead: It would be excluded from the ranking.
Auditor: Can you show me where the exclusion happens?
Contributor (navigating the scorer): It looks like incomplete runs get a partial score. The missing metrics default to the median of the completed runs.
Auditor: So an early-terminated run can rank above completed runs on metrics it never produced. Has that happened?
Lead: I would have to check. I don't know.
Rubric application. On Architecture and design the team passes: descriptions were detailed, the composite-versus-gates decision was justified, and dataset and optimization questions (not excerpted) were answered correctly. On Software implementation the team fails: the promotion path is the safety-relevant path of this artifact, the answer given contradicted the code, and no attendee could establish whether an unintended promotion had occurred. Under the Appendix A implementation rubric this is a significant gap in explaining how the code functions combined with a meaningfully incorrect answer, which grades as Failing regardless of the passing architecture dimension.
Outcome. The contribution receives a failing grade. Promotion from the harness is gated pending remediation: the team must determine whether any early-terminated run was promoted, correct the default-scoring behavior, and demonstrate understanding at a re-audit. The case illustrates two design points from the body. First, single-hop questions were insufficient. The gap surfaced on the second hop of a follow-up, which is why examinations use adaptive follow-up questioning rather than a fixed questionnaire. Second, every process metric was green while the gap accumulated, which illlustrates the fact that healthy metrics cannot certify comprehension.
Case B: a deferred anomaly that grades as a pass
Setup. The audited artifact is the evaluation report accompanying a release candidate, prepared by a three-person team. During preparation the auditors noticed in the run logs that one benchmark suite had been executed twice. The engineer responsible was required to attend by the team lead.
Excerpted exchange (condensed).
Auditor: Your run log shows this suite ran twice, and the first run shows a six-point swing on one reasoning benchmark. What happened?
Engineer: The first run had a variance spike concentrated in one benchmark shard. Three causes were plausible: a data-loader ordering bug, resource contention on the shared cluster, or true run-to-run variance. I re-ran with a fresh seed on a reserved node, got results within the historical band, and shipped the report.
Auditor: Which cause was it?
Engineer: I don't know. I logged it as unresolved. If it was loader ordering it would only affect shard seven, and the shard-level scores in the second run were within a point of historical values, so the reported numbers don't depend on the answer. If it recurs, the loader is the first place to look.
Auditor: Why defer rather than root-cause?
Engineer: The release gate needed the report that week. Root-causing meant reproducing the spike on the shared cluster, roughly two days. I judged that the bounded impact didn't justify the delay, my lead agreed, and the log records that decision.
Auditor: Did you record the anomaly in the bug tracking system?
Engineer: Let me check (pulls up bug tracking system and searches). No I guess I didn't.
Rubric application. Under Key findings and Analysis of results this grades as Passing: the anomaly was identified, plausible causes were enumerated, the impact on the shipped claim was bounded correctly, and the deferral decision was recorded and defensible. There is no knowledge gap. The engineer understands the system well enough to state precisely what she does not know and why it does not affect the reported results.
Outcome. Pass. The audit report carries a recommendation that the team improves processes to track unresolved issues to be triaged with explicit ownership, so that deferred anomalies are revisited rather than forgotten. This case illustrates how comprehension failures are narrower than other kinds of process shortcomings: properly tracking and following up on an unresolved question that is honestly disclosed and correctly bounded is a diligence matter, handled through recommendations and coaching. The audit fails teams that cannot explain what they built. It does not fail teams for disclosing what they chose to defer or for a lack of diligence in investigation. This distinction is important so there's a clean signal of losing understanding (whereas any R&D process will always have many aspects that might be improved). It also helps improve honesty, by removing incentives to hide issues.
Hi Ronald! From just the description, I can't tell what these audits will be like. Do you have a worked example?