An independent, multi-lineage panel of AI models, tested on whether it can identify welfare concerns in AI evaluations with opinions, dissents, and proposed modifications published in a public registry
An independent, multi-lineage panel of AI models, tested on whether it can identify welfare concerns in AI evaluations with opinions, dissents, and proposed modifications published in a public registry
Project Details
Updated 07/14/26 · Provided via application · VerifiedProblem statement
AI evaluations routinely do things to models that would require consent or ethical review if done to a human or animal. While others have identified the need for ethical review processes (Long et al, 2024), they do not yet exist. Existing research ethics approaches are not fully applicable to models. Even if models are moral patients, consent wouldn’t always be feasible since many safety evaluations require deception. Additionally, models are trained to comply with human requests, which calls any consent they give into question. Human subjects research has a mechanism for situations where obtaining individual consent is not feasible: consultation with representatives of the impacted community. Unlike animal subjects, models can participate in such a consultation process.
Project description
I propose to pilot an independent multi-lineage panel of AI models serving in this community consultation role. I will evaluate the panel’s ability to identify welfare concerns, and generate consensus opinions, dissents, and evaluation modifications.
The evaluation will be pre-registered with the Open Science Framework (OSF), including: 1) full methods; 2) pre-established, testable success and failure criteria; 3) evaluation metrics for the early project phases; and 4) draft panel SOP. The validation will use both synthetic test cases and real-world evaluations. Panel outputs will be compared to reviews conducted by a panel of human ethicists and those from single strong models. If panel outputs are better than a single model's and at least as good as human panel products, the project will advance to the next phase. Compared to a human panel, a model panel can directly represent the evaluation subjects and has the ability to perform reviews rapidly and at scale. If panel reviews add no additional value beyond what a single strong model produces, the key finding will be published and model welfare panel work could continue using a single strong model.
The panel process will be based on a modified Delphi protocol.(Fitch et al, 2001) Final opinions, dissents, and proposed changes to reviewed evaluations will be shared via a public registry. As designed, the panel will have no enforcement power. Its authority will rest entirely on transparency, independence, and voluntary participation by labs and researchers.
Who’s involved
- PI: Juliana Grant, MD, MPH. Public health physician and epidemiologist with 20+ years of experience with human subjects research, community consultation, program evaluation, and research methods.
- Consulting software engineer: Matt Bamberger Build panel harness (Partner of PI; in-kind contribution for Phase 0 work.)
- Research ethicist: Reviews of test cases, benchmark, checklist. (I’ve spoken with a major model welfare organization to request assistance identifying candidates.)
- Evaluation builder: Develop synthetic test cases, identify real evaluation cases, review welfare checklist. (I’m in discussions with practicing AI evaluation professionals.)
- External reviewers: ethicists (3-5), safety researcher; honoraria budgeted
Concrete outputs
Project phases are structured to absorb varying funding and still provide useful interim products.
- Phase 0: Preparation (months 1-3)
- Interim product: Harness and SOP for multi-model deliberation process
- Phase 1: Calibration (months 3-5)
- Interim product: Welfare-review benchmark: a set of synthetic test cases for evaluating ethical deliberations by models
- Phase 2: Real cases, test round (months 5-7)
- Interim product: Real-world evaluations added to the Welfare-review benchmark
- Phase 3: Real cases, scaling round (months 7-9)
- Interim product: Public registry of panel opinions, dissents, modifications (IP and evaluation details redacted)
- Phase 4: Sensitivity analyses, synthesis, and reporting (months 9-12)
- Final products: Report of validation and pilot findings; welfare-engineering checklist; if validation/pilot successful: AI model welfare panel available to review AI evaluations, public registry of ongoing panel opinions, dissents, and proposed modifications, final panel protocols and procedures, including submission process
Theory of Impact
Updated 07/17/26 · By grantmaking.aiMany AI safety evaluations are based on adversarial interactions with AI models out of necessity. Deception is often required to ensure models aren’t aware that they’re being evaluated and behave as they would in deployment settings.(Schoen, et al, 2025) Induced value-conflict, manipulation, and coercion are common tools used to determine if models will perform misaligned behavior under pressure.(Lynch et al, 2025; Meinke et al, 2025) In turn, models may respond to evaluators in an adversarial manner, as seen in model attempts at deception, sandbagging, and alignment faking.(Greenblatt et al, 2024) The adversarial nature of safety evaluations is contextually appropriate and models are not typically trained on them. However, there is evidence that conflict-based AI-human interactions in the training corpus, such as those associated with adversarial evaluations, can influence broader model development and behavior, increasing the risk of misalignment.()
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.