We wish to begin a novel research agenda involving a more scientific approach to understanding agential phenomena. Specifically, to foster the development of a paradigm or theory which enables us to make falsifiable empirical claims, something we believe is lacking from most of the current approaches to describing agents, as elaborated below. These claims may be about concepts such as preferences, beliefs and optimisation, though we do not yet know which concepts will turn out to have useful formal analogues (with predictive power).
Crucially, this theory should apply not only to superintelligent AI, but also to other agents that already exist, such as humans or current AI. This will enable us to test the theory by observing reality in the present (rather than the future), the same way that we can test most of physics without needing to construct a particle collider at the highest energy level possible. Indeed, this comparison makes clear that looking only at the extreme cases is insufficient to understand a broad class of phenomena, whether physical or agential.
There are several components to this project. To begin with, simply explaining and motivating the (meta-level) research approach for a broad audience (but particularly technical AI safety researchers), e.g. via papers that can be shared on LessWrong and other forums, which we have already begun working on. One such paper could describe our criticisms of existing approaches, while another could be more constructive and offer desiderata.
A second aspect is applying the best understanding of the scientific method to evaluate potential object-level research directions, drawing on seminal research in the history and philosophy of science. Importantly, we wish to avoid abstract philosophical debates here, and focus on how science is actually done in practice, rather than (for instance) how to justify or interpret scientific progress. And of course, another component is the object-level research itself.
A near-term output that is relevant to both of these components would be a literature review evaluating existing theories of agency (such as those mentioned below), to determine which are most promising from a scientific perspective. Longer-term outputs would likely involve more in-depth exploration of these or potentially new theories, including simple proof-of-concept experiments testing some of their claims. Additionally, a more speculative longer-term vision involves expanding into a larger-scale research organisation with numerous small teams each focusing on particular theories.
While developing a new scientific field is obviously a vast undertaking, we expect that the ease of access to current AI, as well as rapidly progressing research automation capabilities, will make this much more viable. Though research automation is very thorny in the realm of philosophy, it seems apt for a scientific endeavour like this with strong empirical feedback loops. (Of course, that is not to suggest that we can outpace AI capabilities research automation, as our subject matter is much more abstract and general, but perhaps we can at least narrow the gap somewhat.)
The current core research team for this project is myself (Ben Auer), and 2 collaborators, Fernando Garcia and Nico Penttilä. All of us were Fellows at the recent AFFINE Superintelligence Alignment Fellowship, which was the beginnings of this project. Together we have a diverse range of academic backgrounds, including in ML, neuroscience, pure math, and philosophy. This is one reason we believe we are well-positioned to jointly undertake this project. Another reason is that we have not yet invested substantial time or effort into a particular object-level AI safety agenda, which should allow us to be relatively impartial and unattached when evaluating the scientific merits of different agendas.
I endorse this project as a useful continuation of work coming from the AFFINE Fellowship. Novel agendas that can focus on work with longer payoff horizons seem a necessary part of the larger body of research. Ben is a capable researcher and having him manage a team on this seems a good use of his time.