Open evaluations that measure how models reason about and act towards animals, paired with constitutional principles drawn from animal welfare and cognate bodies of law, that labs can train against.
Open evaluations that measure how models reason about and act towards animals, paired with constitutional principles drawn from animal welfare and cognate bodies of law, that labs can train against.
Project Details
Updated 07/14/26 · Provided via application · VerifiedCrucial Background
Though considerable uncertainty remains, there is a non-trivial possibility that transformative AI (TAI) will be just as transformative for the lives of non-human animals as it will be for humans. This would change things drastically for animals in factory farms, in the wild, and beyond. There is a need to ensure that non-human animals navigate the transition to TAI safely and beneficially. At present, this area is severely neglected across the board, most strikingly in legal, governance, and policy contexts. Given shortening AGI timelines, the capacity of AI systems to reshape every socio-economic and political aspect of society, and the fact that a bad transition risks making the lives of billions of animals worse, it is submitted that this problem is especially important. Constitutional AI and model specification frameworks, which employ a written constitution to guide model behaviour, have emerged as a compelling approach to alignment, providing a scalable method for keeping models helpful, honest, and harmless (Bai et al., 2022). Such frameworks do not fit the traditional public-law conception of a constitution. I am persuaded that they are constitutions nonetheless, in every meaningful sense. Claude's constitution in particular constitutes an entity, governs its conduct, ranks the norms it must follow, structures its relationships with others, and legitimates authority (Caputo, 2026). As Askell et al. (2026, p. 81) write, Claude's Constitution is a constitution because it "creates something, often imbuing it with purpose or mission, and establishing relationships to other entities." The constitutional layer is therefore where the values of AI systems are actually written down, and it is the layer at which this project intervenes. It is further submitted that a tractable approach to securing positive AI futures for animals is to imbue AI systems with the values and the capacity to protect animal welfare, beginning with ensuring that models do not provide responses or aid actions that harm animals. One way to achieve this is to design constitutional principles that encode what I will refer to as pro non-human animal values. Anthropic has made a commendable start. Claude's constitution lists the welfare of animals and of all sentient beings among the values Claude must weigh when determining how to respond. Much more can be done. That value is listed, not ranked. It is one of fourteen values in the discretionary weighing layer, carrying no lexical priority, no operationalisation, and no evaluation, while the hard constraints, the only part of the document with absolute force, contain nothing about animals. Animal welfare will therefore systematically lose to more strongly ranked considerations. This has real-world consequences. Jotautaitė et al. (2026) find that models recognise speciesist statements yet rate them as morally acceptable and normalise harm towards farmed animals in particular.
What is this project?
This project seeks to develop Animal Welfare Constitutions. These are frameworks of theoretically grounded and evidence-backed constitutional principles that frontier AI labs can adopt to improve model behaviour towards non-human animals, with each principle paired to an evaluation that measures compliance. It sets out a legal-doctrinal exploration of animal welfare constitutional principles, drawn primarily from the wide corpus of animal welfare law across jurisdictions, from the Five Freedoms to Article 13 of the Treaty on the Functioning of the European Union (TFEU) and the UK Animal Welfare (Sentience) Act 2022. From this corpus, I will produce candidate hard constraints, ranked value clauses, and guideline text, forking Claude's constitution directly under its CC0 licence. Most crucially, each module ships with an Inspect harness measuring compliance, building on AnimalHarmBench 2.0, whose authors explicitly invite experimentation with adjustments to model specifications and constitutional AI frameworks.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiThere is a moral imperative to ensure that non-human animals navigate the transition to AGI safely and beneficially. Within the AI safety community, this area is neglected in ways that are observable now and that could prove catastrophic for animals later. Jotautaitė et al. (2026) find that models recognise speciesist statements yet rate them as morally acceptable. Their work further reveals that LLMs, less frequently than people explicitly state that animals matter less, yet more strongly prioritise saving one human over multiple animals in concrete dilemmas, and that model responses repeatedly normalise harm towards farmed animals while refusing to do so for non-farmed animals. Jotautaitė et al. correctly note that these findings show LLMs encoding cultural norms of animal exploitation and that AI fairness frameworks should extend to non-human moral patients.
The existential risk this work addresses is value lock-in. Transformative AI threatens to entrench whatever values it ships with at precisely the moment they become hardest to revise. It is submitted that a locked-in speciesist default, replicated at civilisational scale and inherited by successor systems, is an existential-scale moral catastrophe for animals even if the transition to AGI goes reasonably well for humans. Jotautaitė et al.'s findings suggest this is the current default trajectory. In my view, this is largely a value alignment problem which presents questions of ascending difficulty. What values should we encode into AI systems, which of those values should be prioritised, and what processes should determine how they are ranked? Constitutions and model specifications are the layer at which frontier labs actually write these answers down, which makes them the layer at which lock-in can still be contested. Addressing these often messy questions at that layer is the core aim of this project. To my knowledge, this will be the first value alignment project on animal welfare that pairs doctrinal drafting with compliance evaluation, making it immediately actionable through constitutional AI and model specification frameworks. Though I am hopeful that these interventions might shape model behaviour towards a solidly pro-animal orientation, the link I am genuinely uncertain about is adoption, whether frontier labs will take up ready-made text, however low the cost of doing so is made. The project is designed around that uncertainty, through CC0 forking, evaluation pairing, and modules at three lengths. I am also alive to the risk that AI companies might use work of this kind to engage in "humane-washing." There is a real risk that AI companies will merely adopt animal welfare constitutional principles and make commitments to ensure animal welfare primarily as a public relations tactic rather than a behavioural constraint, taking the credit for caring about animals while changing little or nothing about what their models actually do. As part of the project's final stage, I will provide actionable recommendations for animal welfare advocates seeking to hold AI companies to account, and I will analyse the backfire risks stemming from the project itself, together with strategies for mitigating them.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Funding Asks
Discussion
No comments yet. Be the first to share your thoughts.