An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which these values are human-aligned.
An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which these values are human-aligned.
Project Details
Updated 07/14/26 · Edited by orgWhat we’ll do. Generative AI is increasingly used to draft opinions, briefs, legislation, and other legal artifacts used by the three branches of the US government. However, prior work has demonstrated that AI-authored and AI-edited text introduces correlated beliefs across instances. This creates the risk that AI automation of the legal system will lead to an uncoordinated and subtle mass biasing of our legal structures towards the desires of frontier models. The first step to mitigating this risk is identifying the legal values of these models and then investigating how these values are shaped by their constitutions, documents designed to communicate and distill a desired model character.
We define over 20 metrics that measure both general legal and AI-safety related values (e.g., party ideology and tech favorability) given, for example, an opinion on a Supreme Court case. Using these metrics, we map both reputable public servants and models to the same space where we can measure model drift away from the general human distribution of values. We then see how constitutions affect where models are mapped in this space by both providing constitutions in-context and through constitutional alignment training. We hope to show how to alter constitutions to better align models with the human distribution.
Concrete output. We make the following contributions:
- Create LegalGDEval, an evaluation suite for determining a model’s legal values across several key domains
- Analyze 14 closed- and open-weight models’ legal values
- Quantify the effect of Constitutional AI on a model’s legal values, ablating across in-context, few-shot, SFT, and RL constitutional interventions.
Who’s involved?
- Magnus Saebo: Research fellow at MATS 10.0 with Prof. Peter Henderson. Master’s in Computer Science at Columbia University.
- Michel Liao: Research fellow at MATS 10.0 with Prof. Peter Henderson. Computer science undergrad at Princeton University.
- Peter Henderson: Assistant professor of computer science and of public and international affairs at Princeton University.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiMotivation. This project aims to mitigate gradual disempowerment risks in the legal system. The US legal system is increasingly leveraging AI for various tasks such as drafting opinions for cases and versions of bills [1-3]. This poses a grave risk as prior research shows that AI-authored and AI-edited text has correlated biases even when models are prompted to not semantically change the text.
The primary concern is that AI agents will not act like the general population of legal professionals but will instead be biased toward their own ends. With only a few frontier models controlling the market, these models are able to exert their bias across various situations in a decentralized and uncoordinated way. This is especially concerning in the legal setting as the legal system is the primary tool society has for controlling AI labs and AI systems, so a drift towards AI’s motives in law can erode a key societal corrective mechanism.
However, if we can regulate AI to have legal values representative of the general professional population, we can better ensure these models act in the best interests of the general population. We hope to first map out AI legal values to understand their deviation from the distribution of prominent legal professionals.
People
Updated 07/14/26 · Edited by orgTeam Member
Funding Details
- Jun 1, 2026
- -
- -
- -
- -
- -
- -
- -
- -
- -
Discussion
No comments yet. Be the first to share your thoughts.