I am requesting funding to transition into full-time x-risk research, specifically by starting an institute focused on embodied AI and the potential for misaligned intelligent systems to enter the physical world. The institute consists of me and my co-founder Sara Filipcic, who has a background in entrepreneurship spanning digital wellbeing and the impact of technology on humanity. My background spans affective neuroscience, computational psychology, behavioral experimentation, and LLM safety research. I recently completed my PhD in Brain & Cognitive Sciences, studying the failures and vulnerabilities of complex intelligent systems in pursuit of grounding x-risk research in established theories. I have published research investigating the neural basis of empathy (https://direct.mit.edu/imag/article/doi/10.1162/imag_a_00110/119822), moral contexts that facilitate the spread of misinformation (https://psycnet.apa.org/record/2025-66798-001), and, most relevant and recent, model organisms of misalignment inspired by psychological dark traits (https://arxiv.org/abs/2603.06816).
The goal of the institute is to pursue rigorous research in order to build regulations and governance around how systems are deployed in the physical world. The research would be led by me, beginning with LLM research that extends directions I have laid out during my PhD. Concrete research outputs include novel evaluations of empathic dissociations in frontier models that mirror human antisocial traits as well as interpretability work that identifies how misaligned behaviors are dissociable within model internals (some of which I have started, here: https://arxiv.org/abs/2605.09773). I recently completed ARENA 8.0, which provided me with concrete skills to pursue these research outputs. Sara will lead the business and policy side of things for the institute, for which concrete outputs include design guidelines for how to build physical AI in a way that is aligned, prosocial, and ethical.