grantmaking.ai Launch Round
This project aims to build an investigations toolkit focused on misaligned agentic AI incidents. The wider project I would like to work on is about establishing investigations as a sub-field of AI Safety and creating an AI Investigations Lab, which could be used to investigate AI incidents, collect evidence, analyze it, correctly identify root causes and ultimately feed the lessons learned back into AI systems to improve their overall safety and security. This would also help improve preparedness for future incident response. Creating an AI investigations lab is incredibly ambitious and requires significant resources, therefore I propose splitting this into smaller and therefore more achievable goals. One of these is developing an investigations toolkit which can be used to review incidents. The toolkit would consist of an investigative methodology, the requirements for evidence preservation, and a playbook that could be used in operations.
AI incidents are not currently being investigated in a consistent way that produces across-the-board root-cause analysis and lessons learned. While risk management in AI Safety and Security is still emerging as a whole, other main areas such as thresholds, capability and risk evaluations, safeguards, and even evaluations of the safeguards themselves currently receive more attention, and rightly so as those layers had to naturally develop first. Meanwhile, the sub-field of incident investigations remains largely unexplored and aside from incident reporting, there is far less research and fewer practical resources available. In my view, this is mainly due to three factors:
1. No known incident with large-scale impact has yet occurred and the risk from an epistemic perspective of not investigating AI incidents is not yet clear to most stakeholders. However, various known incidents do exist today and as AI systems continue to become part of society at every layer, investigative capacity will likely be needed. We already know that certain emerging regulations have started requiring it, but without guidance on how to do this.
2. From a technical and data sensitivity perspective, it is still very challenging to conduct end-to-end investigations of an incident. This is precisely why developing a methodology that includes minimum evidence requirements, built and tested in a controlled environment, would make the difference in future real cases.
3. No standard post-incident investigation methodology has yet emerged. Most investigative work today takes place behind closed doors, mainly in frontier AI labs. There is no visibility externally into how this is conducted and the extent to which root causes are being identified or lessons learned are being developed and integrated back into systems. Yet the need to meaningfully address the residual risk that other controls cannot fully mitigate remains.
I first started researching this idea while completing an ML4Good Technical AI Bootcamp in April 2026. This research became a gap analysis paper, which found that no dedicated investigations methodology currently exists for AI incidents. My focus was specifically on misaligned agentic incidents, starting from examples such as the research published by Anthropic on Agentic Misalignment: How LLMs could be insider threats (June 2025). The paper mapped this gap and argued for the development of a dedicated methodology, adapted from intentional insider threat investigation, before more serious real-world cases emerge.
The investigative methodology I am proposing would be built by adapting investigative practices from other fields and would be validated by producing investigations on publicly available cases, as well as by testing on controlled incidents in a sandbox environment.
In terms of who would be involved, I see this project as an opportunity for collaboration, bringing the AI safety community and subject matter experts from security, risk management, and investigations into the same room. My focus would be on identifying the right experts across the areas needed to advance investigations in AI safety. Regarding my own background, I bring nearly ten years of experience building and scaling security, investigations, and risk management programs across non-profit, international law enforcement, and tech environments, including in orgs such as the WJC, INTERPOL, and Google. I have led cross-functional P0 incident investigations and became part of an AI red-teaming project while at Google, which ultimately led me to further up-skill and my current transition into AI safety. Investigating AI incidents is the area where I believe my skills can make the most impact.