Naturalizing theoretical alignment by initiating the development of a scientific theory that is capable of making falsifiable empirical claims about agents in general, including humans, AGI, and ASI, and thereby about alignment.
Naturalizing theoretical alignment by initiating the development of a scientific theory that is capable of making falsifiable empirical claims about agents in general, including humans, AGI, and ASI, and thereby about alignment.
Project Details
Updated 07/14/26 · Provided via application · VerifiedWe wish to begin a novel research agenda involving a more scientific approach to understanding agential phenomena. Specifically, to foster the development of a paradigm or theory which enables us to make falsifiable empirical claims, something we believe is lacking from most of the current approaches to describing agents, as elaborated below. These claims may be about concepts such as preferences, beliefs and optimisation, though we do not yet know which concepts will turn out to have useful formal analogues (with predictive power).
Crucially, this theory should apply not only to superintelligent AI, but also to other agents that already exist, such as humans or current AI. This will enable us to test the theory by observing reality in the present (rather than the future), the same way that we can test most of physics without needing to construct a particle collider at the highest energy level possible. Indeed, this comparison makes clear that looking only at the extreme cases is insufficient to understand a broad class of phenomena, whether physical or agential.
There are several components to this project. To begin with, simply explaining and motivating the (meta-level) research approach for a broad audience (but particularly technical AI safety researchers), e.g. via papers that can be shared on LessWrong and other forums, which we have already begun working on. One such paper could describe our criticisms of existing approaches, while another could be more constructive and offer desiderata.
A second aspect is applying the best understanding of the scientific method to evaluate potential object-level research directions, drawing on seminal research in the history and philosophy of science. Importantly, we wish to avoid abstract philosophical debates here, and focus on how science is actually done in practice, rather than (for instance) how to justify or interpret scientific progress. And of course, another component is the object-level research itself.
A near-term output that is relevant to both of these components would be a literature review evaluating existing theories of agency (such as those mentioned below), to determine which are most promising from a scientific perspective. Longer-term outputs would likely involve more in-depth exploration of these or potentially new theories, including simple proof-of-concept experiments testing some of their claims. Additionally, a more speculative longer-term vision involves expanding into a larger-scale research organisation with numerous small teams each focusing on particular theories.
While developing a new scientific field is obviously a vast undertaking, we expect that the ease of access to current AI, as well as rapidly progressing research automation capabilities, will make this much more viable. Though research automation is very thorny in the realm of philosophy, it seems apt for a scientific endeavour like this with strong empirical feedback loops. (Of course, that is not to suggest that we can outpace AI capabilities research automation, as our subject matter is much more abstract and general, but perhaps we can at least narrow the gap somewhat.)
The current core research team for this project is myself (Ben Auer), and 2 collaborators, Fernando Garcia and Nico Penttilä. All of us were Fellows at the recent AFFINE Superintelligence Alignment Fellowship, which was the beginnings of this project. Together we have a diverse range of academic backgrounds, including in ML, neuroscience, pure math, and philosophy. This is one reason we believe we are well-positioned to jointly undertake this project. Another reason is that we have not yet invested substantial time or effort into a particular object-level AI safety agenda, which should allow us to be relatively impartial and unattached when evaluating the scientific merits of different agendas.
Theory of Impact
Updated 08/17/26 · By grantmaking.aiAlignment between AI and humans seems to be a key ingredient that would drastically curtail x-risk from AI, at least by bringing it much closer to the default x-risk level already entailed by human society (modulo advanced AI). Specifically, we require robust alignment that is maintained even under pressures such as high levels of intelligence and the ability to self-improve, which threaten most existing approaches to alignment.
A powerful, accurate theory of what agents are has long been coveted in the agent foundations research community. It is widely believed that such a theory would endow us with the understanding needed to engineer robust alignment. While we are not certain whether reality admits of such a theory, we believe that it is a worthwhile endeavor to pursue.
The premise of this project is that a major flaw in most agent foundations research until this point has been a lack of scientific rigor, and that this has prevented the development of a powerful theory, as well as blunted progress on reducing our confusion about agents in general.
More specifically, the current dominant paradigm of Expected Utility Theory (EUT), the basis of decision theory and game theory, is favoured mainly due to the contested normative assumptions it rests on (e.g. choice-wise inexploitability). While there is an aspect of normativity to alignment, there are also central empirical questions, such as which types of objectives generalise well, which types of agents cooperate well, and where mesa-optimisers arise. These questions can and should be separated entirely from the normative aspects, as in other fields of science.
People
Updated 07/15/26 · Edited by orgTeam Member
Discussion
I was a mentor at AFFINE and I interacted with Nico, Fernando and Ben quite a lot during that time period.
What stands out to me is the combination: each of them scores high on intelligence, conscientiousness and openness (imo) at once — high enough on all three that they work well as a team and can make real progress independently. That full profile is rare, and it's exactly what you want in the founding team of a research organisation. When these three decided to team up, it made immediate sense to me.
Whilst I haven't dug into their proposal in detail, I've got a good grasp of who they are individually.
Nico is conscientious, asks good questions, iterates well on feedback, and pivots quickly when he's wrong. He understands science and mathematics deeply, and has had a long-standing background interest in both. If I left him with six months to work on a specified project, I'd expect to see good results when I checked back.
Fernando knows category theory (which clearly means he's smart) and a good deal of mathematics besides, and as a former teacher he communicates well. In a research environment he's connective glue — understanding what different people are working on and sharing it across the group, powered by an intense curiosity of the productive kind. Pointed at the right problems, he'd be a valuable addition to any agent foundations environment.
Ben is a very solid, conscientious researcher who came to this work from the impact side. He works hard, digs for the underlying structure of things, and has spent real time in the agent foundations context. He's the kind of person I'd expect to just keep producing continual good work.
If I had to name the factor that will determine whether this works, I think it's iteration speed: how often they can get their thinking out in public, and how quickly they learn from the response. That's what I'd watch. Concretely, I'd fund the $50k and see whether a couple of genuinely interesting posts come out of the first six months. At that price it's a bet well worth taking and if it pays off, this is a team worth continuing to fund and bet on.
I endorse this project as a useful continuation of work coming from the AFFINE Fellowship. Novel agendas that can focus on work with longer payoff horizons seem a necessary part of the larger body of research. Ben is a capable researcher and having him manage a team on this seems a good use of his time.