Building an improved open-source suite to understand multimodal models detect misalignment, for the technically minded community, and communication with this group of audiences
Building an improved open-source suite to understand multimodal models detect misalignment, for the technically minded community, and communication with this group of audiences
Project Details
Updated 07/18/26 · Edited by orgSince 2025, I have been working on understanding multimodal models by running experiments and looking at how to detect misalignment. Misalignment in AI refers to artificial intelligence systems that fail to align with human values, goals, or intentions. My experiments have shown some results.
What are these results? One experiment on the COCO dataset (Common Object in Context dataset by Microsoft) reveals critical gaps between true semantic understanding and benchmarking. True semantic understanding refers to a genuine grasp of meaning, and benchmarking means comparing against a standard. It shows 0.216 average cosine similarity (a measure between two vectors) for matching text-image pairs versus near-zero for mismatching ones.
Besides, another experiment involves a compositional copying (giving a model a text cue and seeing what visual feature it copies) task using the Visual Genome (a dataset containing relationships between objects and scenes) dataset. It shows confident predictions on brittle features unlikely to survive distribution shifts (data that shift away from the training data) in the real world. an 86.8% probability score (on a scale of 10, its confidence level is roughly 9). However, it relies on tiny embedding (numerical vector representing objects) differences (0.019 cosine similarity and 0.074 patch-level probability).
I have built an open-source suite for the community to verify and extend my research.
What I will continue doing is:
· Designing experiments to understand multimodal models using various techniques, including mechanistic interpretability and representation engineering
· Improving on the open-source suite to help the community verify and extend what I have done
· Building tools to detect misalignment, including undesirable behaviours, for current and future models
Who is involved: I will take the lead in completing the technical work and sharing it. I will hire 1 to 4 AI safety researchers to work with me, if possible. I might need to find the right kinds of collaborators later. The entire community will have the opportunity to get involved by sharing and giving constructive feedback
The concrete output is an improved open-source suite published on GitHub for the technically minded community, communication with this group of audiences, and perhaps papers or reports down the line.
Theory of Impact
Updated 07/18/26 · By grantmaking.aiMy work can reduce x-risk because it helps others understand what multimodal models are like and detect misalignment in models. 4+ researchers and data scientists have already investigated the open-source suite, which I have published on GitHub. Here are the reasons justifying my approach.
To see how my work is relevant, we define what the theory of impact means and what x-risk refers to. From my understanding, the theory of impact, which is derived from the theory of change, describes how we get from our inputs to the desired impact for change. X-risk refers to existential risk. Based on how Nick Bostrom defines it in 2002, it is a risk that can cause the early extinction of intelligent life on Earth. Besides, it may lead to the permanent and significant destruction of its potential for the future development we want.
First, as we have observed over the last few years, AI has moved so quickly, and we do not have reliable ways of predicting its real-world performance. The International AI Safety Report 2026 states that general-purpose AI capabilities have continued to improve, even when their capabilities are jagged, meaning that they excel in some tasks while failing at some other simple tasks. Benchmarks often fail to predict real-world performance since many models have been trained using data from these same benchmarks (data contamination). It leads to inflated scores that do not reflect a model’s genuine ability. We need radically different approaches to understanding these models and their capabilities reliably.
People
Updated 07/11/26 · Edited by orgTeam Member
Funding Details
- Aug 1, 2026
- Jan 1, 2027
- 6 months
- -
- -
- -
- -
- -
- Seeking first grant
- -
Discussion
No comments yet. Be the first to share your thoughts.