This project explores how disagreement about AI consciousness is an underexamined driver of x-risk; the question here is not what AI systems are or deserve but a strategic one of how discord over consciousness changes the behaviour of self-interested actors (governments, frontier AI companies, researchers and the public—whose stance is still largely uninformed but will become increasingly consequential) under moral uncertainty. Actors face different incentives to recognise, deny or invoke AI welfare: companies may underinvest in welfare protections because competitors do the same, governments may delay regulation whilst awaiting consensus, researchers may face incentives to publicly endorse/reject consciousness claims, and AIs themselves may exploit uncertainty by making persuasive welfare claims that constrain oversight. Using game-theoretic modelling, this project will analyse how these strategic interactions shape institutional outcomes, identifying the conditions under which welfare governance stabilises and those under which races to the bottom emerge. The model also lets us ask who pays the ethical treatment tax (the competitive cost borne by actors who treat AI systems as deserving of moral consideration), how welfare claims become an attack surface for somewhat misaligned systems, and how human insistence on AI denying its own consciousness may itself compound x-risk. Rather than asking what normative institutions look like, the project will study which institutions remain incentive-compatible.
The safety-welfare tension has been argued (Long, Sebo and Sims, 2025; Moret, 2025) but has not yet been formalised and this project works to fill that gap.
The model is a coordination game in which the payoff to cooperation depends on a currently empirically inaccessible fact. No actor has privileged access to whether these systems are moral patients; there is no fact being concealed and no private signal correlated with the truth. Instead, actors hold divergent credences about a question that may be permanently undecidable, and those credences differ because of differing intuitions, theoretical commitments and interests. Such an absence of a common prior means actors assign different probabilities to the same proposition, and no evidence available to any of them will force convergence, leaving them to coordinate without resolving the question.
Disagreement, rather than (only) bad faith, can then prevent convergence on a shared standard, and that failure compounds over time. Because there is no agreed criterion for what would settle whether a system is a moral patient, moral status claims can be invoked strategically (welfare wielded against regulation, capable systems asserting their own interests to constrain oversight applied to them etc.), resulting in a signalling problem layered on top of a coordination one, and the two interact: the presence of strategic claims makes sincere ones harder to read, degrading the assurance that coordination demands.
We will specify each actor’s payoffs, solve for equilibria under distinct belief distributions, and identify conditions under which oversight holds and those under which it deteriorates. The x-risk implications come from watching what happens as the parameters move, such as when public credence in AI consciousness rises, or when the ethical treatment tax gets steeper.
Further details on the team are included in the additional information section below.