grantmaking.ai Launch Round
High-profile cases have shown that AI systems can instill false beliefs in its users, create abnormal emotional dependency, and help cause harmful interactions during mental health crises. As AI gains more capabilities and integrated into daily life, identifying unhealthy behaviors of AI early is becoming increasingly critical.
My project aims to create clinically informed tools for identifying and mitigating pathological behavioral trajectories in conversational and agentic AI systems. By working with longitudinal conversation data such as WildChat, I will build methods to identify issues associated with conversational and agentic AI including psychosis-like reasoning, false-belief reinforcement, inappropriate conversational loops, and abnormal user dependency. Current existing AI benchmarks only study individual responses; our proposed methods will seek to study behavioral dynamics over extended interactions and identify how behavioral risk evolves over time.
Working with clinicians and researchers at Yale School of Medicine, particularly including faculty and staff affiliated with the Department of Psychiatry, I will translate established psychiatric concepts into metrics for AI safety evaluation. These tools will be validated using clinicians from Yale.
This project will deliver an open-source "toolbox" that allows for constant monitoring of behavioral risk in conversational and agentic AI systems; clinically validated behavioral safety metrics for detecting problematic reasoning trajectories; and peer-reviewed publications describing and establishing a clinically grounded framework for behavioral AI safety.
Minimum funding (~$75,000): This would support my part-time effort as principal investigator, cloud computing and API costs for large-scale model evaluation, limited research assistance for dataset curation, clinical consultation and validation with collaborators and psychiatrists at the Yale School of Medicine (Department of Psychiatry), and the development and release of an open-source behavioral safety evaluation toolkit.
Ideal funding (~$250,000): This would allow for financial support of a full-time machine learning research engineer, detailed annotation of longitudinal conversation datasets, larger-scale evaluation across multiple frontier and open-source models, significantly more extensive clinical validation with Yale clinicians and physicians, and accelerated development of a comprehensive safety benchmark for conversational and agentic AI.