grantmaking.ai Launch Round
Building an improved open-source suite to understand multimodal models detect misalignment, for the research community
Since 2025, I have been working on understanding multimodal models by running experiments and looking at how to detect misalignment. Misalignment in AI refers to artificial intelligence systems that fail to align with human values, goals, or intentions. My experiments have shown some results.
I have built an open-source suite for the community to verify and extend my research.
What I will continue doing is:
· Designing experiments to understand multimodal models using various techniques, including mechanistic interpretability and representation engineering
· Improving on the open-source suite to help the community verify and extend what I have done
· Building tools to detect misalignment, including undesirable behaviours, for current and future models
Who is involved: I will take the lead of completing the technical work and sharing it. I will hire 1 to 4 AI safety researchers to work with me, if possible. I might need to find right kinds of collaborators later. The entire community will have the opportunity to get involved by sharing and giving constructive feedback
The concrete output is an improved open-source suite published on GitHub, communication with non-technical audiences, and perhaps papers or reports down the line.
From my understanding, theory of impact, which is derived from theory of change, describes how we get from our inputs to the desired impact for change. X-risk refers to existential risk. Based on how Nick Bostrom defines it in 2002, It is a risk that can cause early extinction of intelligent lives on Earth. Besides, it may lead to the permanent and significant destruction of its potential for the future development we want.
My work can reduce x-risk because it helps others understand what multimodal models are like and detect misalignment in models. 4+ researchers and data scientists have already investigated the open-source suite which I have published on GitHub. Here are the reasons justifying my approach.
First, as we have observed over the last few years, AI has moved so quickly. The International AI Safety Report 2026 states that general-purpose AI capabilities have continued to improve, even when their capabilities are jagged, meaning that they excel in some tasks while failing at some other simple tasks. Benchmarks often fail to predict real-world performance since many models have been trained using data from these same benchmarks (data contaminations). It leads to inflated scores that do not reflect a model’s genuine ability. We need radically different approaches to understanding these models and their capabilities reliably.
Furthermore, our understanding of models and capabilities to detect misalignment in models, especially multimodal models lag. Balasubramanian et al (2025) states that mechanistic interpretability lags behind for multimodal models although multimodal models are the future. On the other hand, Dan Hendrycks and Laura Hiscott (2025) in their article warn us about high investments in mechanistic interpretability without corresponding returns. Their central arguments advise against investing too much in ideas unlikely to work, potential to neglect of more effective ones. We should be more sceptical of give mechanistic interpretability too many resources at the expense of other types of AI safety research.
In my research, I can see that mechanistic interpretability can work in some experiments with certain changes that I have applied to. Applying mechanistic interpretability in its current approach without looking at the bigger picture, will not work well. The current approach usually involves studying a small circuit (a subgraph of neural networks) with a toy model (a small, constructed model). Therefore, we need to combine different approaches, including representation engineering, holistically.
In conclusion, my work directly contributes to understanding of models and capabilities, with experiments to deepen our knowledge of multimodal models and how to detect misbehaviours and misalignment in multimodal models. It is a small, meaningful contribution to preventing existentially risk from AI. It is a different approach from previous research, signifying a shift away from the well-trodden paths. According to the universal hypothesis, all models will likely to converge to similar behaviours. It means that the results from my work will likely be useful for our endeavours with future powerful models.
Paying myself and hiring 1 to 4 AI safety researchers to do the work.