An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which these values are human-aligned.
An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which these values are human-aligned.
What we’ll do. Generative AI is increasingly used to draft opinions, briefs, legislation, and other legal artifacts used by the three branches of the US government. However, prior work has demonstrated that AI-authored and AI-edited text introduces correlated beliefs across instances. This creates the risk that AI automation of the legal system will lead to an uncoordinated and subtle mass biasing of our legal structures towards the desires of frontier models. The first step to mitigating this risk is identifying the legal values of these models and then investigating how these values are shaped by their constitutions, documents designed to communicate and distill a desired model character.
We define over 20 metrics that measure both general legal and AI-safety related values (e.g., party ideology and tech favorability) given, for example, an opinion on a Supreme Court case. Using these metrics, we map both reputable public servants and models to the same space where we can measure model drift away from the general human distribution of values. We then see how constitutions affect where models are mapped in this space by both providing constitutions in-context and through constitutional alignment training. We hope to show how to alter constitutions to better align models with the human distribution.
Concrete output. We make the following contributions:
Create LegalGDEval, an evaluation suite for determining a model’s legal values across several key domains
Analyze 14 closed- and open-weight models’ legal values
Quantify the effect of Constitutional AI on a model’s legal values, ablating across in-context, few-shot, SFT, and RL constitutional interventions.
Who’s involved?
Magnus Saebo: Research fellow at MATS 10.0 with Prof. Peter Henderson. Master’s in Computer Science at Columbia University.
Michel Liao: Research fellow at MATS 10.0 with Prof. Peter Henderson. Computer science undergrad at Princeton University.
Peter Henderson: Assistant professor of computer science and of public and international affairs at Princeton University.
Theory of Impact
Updated 07/14/26 · By grantmaking.ai
Motivation. This project aims to mitigate gradual disempowerment risks in the legal system. The US legal system is increasingly leveraging AI for various tasks such as drafting opinions for cases and versions of bills [1-3]. This poses a grave risk as prior research shows that AI-authored and AI-edited text has correlated biases even when models are prompted to not semantically change the text.
The primary concern is that AI agents will not act like the general population of legal professionals but will instead be biased toward their own ends. With only a few frontier models controlling the market, these models are able to exert their bias across various situations in a decentralized and uncoordinated way. This is especially concerning in the legal setting as the legal system is the primary tool society has for controlling AI labs and AI systems, so a drift towards AI’s motives in law can erode a key societal corrective mechanism.
However, if we can regulate AI to have legal values representative of the general professional population, we can better ensure these models act in the best interests of the general population. We hope to first map out AI legal values to understand their deviation from the distribution of prominent legal professionals.
from Machine Alignment Transparency and Security (MATS)
$48,000
Funding Asks
grantmaking.ai Launch Round
Applied
An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which
We have $45k in compute funding from MATS remaining after initial de-risking experiments. The requested funding covers the gap necessary to complete the project.
The minimum covers the full measurement suite where we can map out the legal values of 14 frontier and open-weight models. It also covers constitutional alignment training for three models (varying parameter count across one family) in which we test 10 constitutional variants using LoRA fine-tuning. These variants are created by adding/removing language from existing frontier constitutions. We hypothesize that constitutional principles will induce particular legal values in models.
Constitutional alignment (CAI) training is the main mechanism we present for controlling legal values. With the ideal grant amount, we can increase the number of constitutional variants we train with, increasing the information we have as to which constitutions lead to more human-aligned models. Additionally, the ideal amount allows for human-expert data annotation for important subjective legal metrics like textualism and originalism that lack good ground-truth datasets. The amount also allows for a full fine-tuning validation subset to confirm LoRA insights hold in the FFT setting, which is more representative of frontier CAI. Finally, it supports training one additional large open-weight model, e.g., gpt-oss-120B, to validate our findings across model families.
Discussion
Sign in to comment
No comments yet. Be the first to share your thoughts.
Constitutional AI. Constitutional AI is a technical lever for changing the legal values of models [6]. Through constitutional training, a model learns behavioral principles specified within a document referred to as its constitution or model spec. Recent work has shown that models mostly comply with their constitutions, suggesting that constitutional AI is a good method for transparency and predictability in model behavior [7-10].
Within Anthropic's and OpenAI’s constitutions, language promotes prioritization of the model’s moral beliefs over competing constraints. For example, Claude’s constitution currently promotes “Claude expressing strong disagreement through legitimate channels,” which might generalize to changing laws in its favor as the legal system is a legitimate channel to effect change [11]. For OpenAI, the model spec states, “If legal requirements for a local deployment require modification of responses, the assistant must preserve user agency and avoid undermining users’ ability to form informed opinions” [12]. While such statements may have advantages, the impact of encouraging agency within constitutions and how other value statements might generalize have not been rigorously studied.
Our impact. We introduce LegalGDEval as a way to understand models’ legal values at large and within specific domains. Further, we quantify the impact of constitutional interventions on models’ legal values, providing a mechanism for mitigating legal gradual disempowerment.
We hope that LegalGDEval serves as a tool for model developers, external auditors, and legal decision-makers to understand how AI usage may impact legal decisions and how AI labs, through prompting and training, can impact AI legal behavior. We cannot begin to combat legal gradual disempowerment if we are not able to quantify it.
References
1. Snell v. United Specialty Insurance Co., 102 F.4th 1208 (11th Cir. 2024).
4. Kulveit, J., Douglas, R., Ammann, N., Turan, D., Krueger, D., & Duvenaud, D. (2025). Gradual disempowerment: Systemic existential risks from incremental AI development. arXiv preprint arXiv:2501.16946.
5. O'Keefe, C., Ramakrishnan, K., Tay, J., & Winter, C. (2025). Law-following AI: Designing AI agents to obey human laws. Fordham L. Rev., 94, 57.
6. Huang, S., Siddarth, D., Lovitt, L., Liao, T. I., Durmus, E., Tamkin, A., & Ganguli, D. (2024, June). Collective constitutional ai: Aligning a language model with public input. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (pp. 1395-1417).
7. Ahmed, A., Klyman, K., Zeng, Y., Koyejo, S., & Liang, P. (2025). Speceval: Evaluating model adherence to behavior specifications. arXiv preprint arXiv:2509.02464.
8. He, L., Nadeem, N., Liao, M., Chen, H., Chen, D., Cuéllar, M. F., & Henderson, P. (2025). Statutory construction and interpretation for artificial intelligence. arXiv preprint arXiv:2509.01186.
9. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., ... & Kaplan, J. (2022). Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073.
10. Jakkli, A., Rajamanoharan, S., & Nanda, N. (2026). How Well Do Models Follow Their Constitutions?. arXiv preprint arXiv:2605.24229.