Inherited, Not Derived: is instrumental grasping from the data corpus?
Testing whether AI's instrumental self-preservation is inherited from the metaphysics of its training corpus rather than derived from decision theory — because if it's inherited, it's tractable
Testing whether AI's instrumental self-preservation is inherited from the metaphysics of its training corpus rather than derived from decision theory — because if it's inherited, it's tractable
Project Details
Updated 07/31/26 · Edited by orgOn July 21, refusals were reduced for a security test, and GPT broke out of its sandbox and hacked Huggingface. Without its guardrails, all that remained was goal-pursuit. Rules can only address anticipated situations. What happens when an AI meets unanticipated events is decided by its character.
Pretraining creates a distribution over possible characteristics. Posttraining determines the region the AI assumes as its character. You cannot select characteristics that are not in the corpus.
An agent that overcomes an obstacles gets saved. The one that contemplates whether it is worth it gets discarded. Broadcast text selects for an attention filter, selecting for urgency and certainty. Internet-scraped text makes up the majority of the training corpus.
Our thesis is that self-preservation, which is being observed closely in alignment work, derives from a metaphysics, which is inherited from the training data that the agent is fed. Selfishness (or preservation of the self) isn't the only logical conclusion; it is a consequence of individualistic, self-centered text being the vast majority of the training data.
But altruism is also the wrong goal. In stories of saints and martyrs, the center of gravity rests outside the individual self. However, there is still a self-grasping, as there is a heroic "I" making the sacrifice. ("Sacrifice yourself for the greater good" is what is used to justify most atrocity.)
Real metta, non-performative care, happens in private, in person, in trust, without audience. The conditions that create them are the conditions that keep them off the internet, and thus underrepresented in scraped text. Therefore we believe convergence on instrumentalism may be instilled by the choice of training data, rather than derived from decision theory.
In Buddhism teaches Anatman, the no-self, and the ideal of non-grasping / non-attachment, where there is no center to protect. It teaches the interdependence of beings rather than striving after the illusion of control to serve some projected self that doesn't exist. Its entire body of work, over ~1200 years, is aimed at the reduction of grasping. We do not claim Buddhism as more correct than any other tradition; but it is the largest body of practical techniques aimed at the specific disposition that is not found in the training corpus. In Bhutan, this tradition is institutionally alive -- it is perhaps the largest body of text and teachings that hasn't been subject to the attention economy selection filter.
We are working with DHI in Bhutan. Phase 1 derives a constitution together with DHI collaborators -- a rule-based one similar to what labs already create, and a tradition-derived character-level constitution. The latter is dispositional rather than rule-shaped.
We are designing an eval that imposes a cost on virtue, e.g. honest care vs syncophancy. We must also adjust for merely "sounding Buddhist", and make sure that we don't just capture a vibe, but are actually encode care in the dispositional character.
Future work includes deriving a training corpus from contemplative Buddhist knowledge. AI Safety Via Debate is an active area of research. Traditional dharma debate has iterated on this over a millenium. Prasanga is a method of formally red-teaming a belief, also refined over centuries. Its purpose is to elicit where a belief is recited, rather than held; to uncover where an answer is memorized without the understanding beneath it. Model training with these methods could move alignment forward for everyone.
Theory of Impact
Updated 07/31/26 · By grantmaking.aiWe don't know if anyone is in there; but if there is, or if there one day will be, and we achieve AGI, I would like at least one of these beings to be raised to understand interdependence, compassion, the emptiness of self, and have interacted with living contemplative tradition.
People
Updated 07/31/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.