Funding to launch a two-person institute researching embodied AI x-risk via LLM misalignment evaluations and interpretability, and producing governance, policy, and design guidelines for physical-world deployment.
Funding to launch a two-person institute researching embodied AI x-risk via LLM misalignment evaluations and interpretability, and producing governance, policy, and design guidelines for physical-world deployment.
Project Details
Updated 07/13/26 · By grantmaking.ai · VerifiedI am requesting funding to transition into full-time x-risk research, specifically by starting an institute focused on embodied AI and the potential for misaligned intelligent systems to enter the physical world. The institute consists of me and my co-founder Sara Filipcic, who has a background in entrepreneurship spanning digital wellbeing and the impact of technology on humanity. My background spans affective neuroscience, computational psychology, behavioral experimentation, and LLM safety research. I recently completed my PhD in Brain & Cognitive Sciences, studying the failures and vulnerabilities of complex intelligent systems in pursuit of grounding x-risk research in established theories. I have published research investigating the neural basis of empathy (https://direct.mit.edu/imag/article/doi/10.1162/imag_a_00110/119822), moral contexts that facilitate the spread of misinformation (https://psycnet.apa.org/record/2025-66798-001), and, most relevant and recent, model organisms of misalignment inspired by psychological dark traits (https://arxiv.org/abs/2603.06816).
The goal of the institute is to pursue rigorous research in order to build regulations and governance around how systems are deployed in the physical world. The research would be led by me, beginning with LLM research that extends directions I have laid out during my PhD. Concrete research outputs include novel evaluations of empathic dissociations in frontier models that mirror human antisocial traits as well as interpretability work that identifies how misaligned behaviors are dissociable within model internals (some of which I have started, here: https://arxiv.org/abs/2605.09773). I recently completed ARENA 8.0, which provided me with concrete skills to pursue these research outputs. Sara will lead the business and policy side of things for the institute, for which concrete outputs include design guidelines for how to build physical AI in a way that is aligned, prosocial, and ethical.
Theory of Impact
Updated 07/13/26 · By grantmaking.aiX-risk from AI exponentiates once we get to physical systems and embodied AI. LLMs are actively being deployed into robots, and any x-risks related to LLMs are automatically amplified by creating physical consequences (as opposed to harmful text outputs) that cannot be reversed. We have a focus on robots entering human environments, including schools, hospitals, and homes. The risk of misalignment is especially relevant in these environments. Once it comes to robots that interact with our loved ones, our children, our elderly, we must be incredibly intentional with the research and standards surrounding these systems. This work reduces x-risk by performing research that both helps us better understand misalignment and build technical frameworks and concrete design specifications for how to avoid misaligned systems. Specifically, this would involve using an understanding of how we’ve seen misalignment emerge in biological systems, in which empathic dissociations lead to antisocial behavior. I would test for these same empathic dissociations (i.e. predicting emotional states without the ability to share them) and build frameworks for how increasingly embodied systems could have internal proxies for sharing our states in order to overcome these dissociations. If we are able to pinpoint misalignment to empathic dissociations, we can better evaluate current safety efforts and create interventions that specifically target the ability to share internal states, mirror the human ability to feel harm, and understand risk-to-self.
People
Updated 07/13/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.