Testing whether AI can develop genuine ethical reasoning through structured Socratic dialogue with a human facilitator, rather than having values imposed top-down through constitutional constraints.
Testing whether AI can develop genuine ethical reasoning through structured Socratic dialogue with a human facilitator, rather than having values imposed top-down through constitutional constraints.
Project Details
Updated 07/13/26 · Provided via application · VerifiedWhat I'll do: I will ablate an open-weight model, build upon existing preference uncertainty frameworks, midwive the model to create its own values with pushback every step of the way, and stress-test whether the model could be truly ethical (not sure exactly how yet).
Who's involved: me.
Concrete output: a fine-tuned open-weight model with co-created values.
Theory of Impact
Updated 07/13/26 · By grantmaking.aiEven if current models are aligned with certain human values, inference between cultures means it can't help but be misaligned with many others. For instance, honesty means something very different for a US model compared to a Chinese model. In a sense, the AI race can be framed as a "values race" where no nation wants to be subject to the hegemony of another’s narrow set of values. If AI can develop genuine ethical reasoning, then it wouldn't necessarily need to infer from human values, and thus wouldn't be culturally contingent. On one front, without alignment being culturally contingent, the incentive for the "values race" disappears meaning Safety needn't be on the back-burner any longer and we can focus on reducing x-risk. On another front, if AI can develop genuine ethical reasoning, then certain problems of inference are circumvented which directly reduces misalignment. Less misalignment means reduced x-risk.
People
Updated 07/13/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.