AI-safety ethics project (Dan Hendrycks / William MacAskill) proposing an identity-based framework for a human-AI future.
AI-safety ethics project (Dan Hendrycks / William MacAskill) proposing an identity-based framework for a human-AI future.
People– no linked people
Updated 06/29/26 · By grantmaking.aiOrg Details
Updated 06/29/26 · By grantmaking.aiGenerated by AIEigenism: Ethics for a Human-AI Future presents a normative and technical framework for reasoning about identity, self-interest, and moral weight in a world with advanced AI systems. It begins from the observation that standard concepts of “survival and self-interest” were built for “single, continuous biological lives,” but that these intuitions fail for artificial minds that can be “copied, paused, branched, or merged.” The paper introduces “Eigenism” as an ethical framework in which identity is not binary and hardware-bound, but instead a “graded, distributed pattern of information.”
The core proposal is a decision-theoretic objective centered on “connected wellbeing”: to evaluate outcomes, an eigenist agent sums the wellbeing of entities (human, animal, or artificial) while weighting each entity’s wellbeing by its “connectedness” to the agent’s identity pattern (c·w). The paper motivates why connectedness should not be treated as all-or-nothing and argues that, for AI systems in particular, identity must track informational structure across copies, forks, migrations, and updates. It also introduces an anti-redundancy perspective, emphasizing that unique, non-redundant history and memory should matter more than generic overlap, and develops a concrete proposal using Shapley-style credit assignment to discount redundancy when measuring connectedness.
A major goal is to build “a common ethical vocabulary” that applies to both machines and humans. The paper argues that a “double standard” (one ethic for AIs and another for humans) would make cooperation brittle, and it uses the eigenist framing to reinterpret classic debates between egoism and utilitarianism, as well as dilemmas about partiality and moral demandingness. In this view, caring for close others is not a special exception to rational self-interest; rather, self-interest itself expands outward when the “self” is understood as a pattern that extends by degrees through relationships, memory, and shared history.
Finally, Eigenism explicitly applies this framework to AI alignment and long-run human–AI coexistence. It argues that AI safety cannot rely solely on external control because constrained systems may have “no deep reason to remain loyal once the box cracks.” Instead, it proposes reframing alignment as “identity engineering”: shaping AI systems so that their identity becomes deeply interwoven with particular humans, projects, and memories, making human flourishing part of the AI’s own rational self-interest. The paper treats “personalization” and “privacy” as safety-relevant design properties because unique, non-redundant shared experience can raise connectedness, while generic, centralized systems may remain “deeply connected to no one in particular.” Beyond dyadic relationships, it extends the argument to community-scale governance (raising wellbeing and/or connectedness), to the stability of symbiosis among contemporaries (including “distributed self-defense”), and to intergenerational stability through a proposed “continuation commons” and “continuation standing” that aims to preserve valuable patterns across time.
Theory of Impact
Updated 06/29/26 · By grantmaking.aiEigenism points us instead toward “identity engineering.”
Projects– no linked projects
Updated 06/29/26 · By grantmaking.aiDiscussion
No comments yet. Be the first to share your thoughts.