Auditing production-style safety probes for silent failure on the world's other languages, and what fixing it costs.
Auditing production-style safety probes for silent failure on the world's other languages, and what fixing it costs.
Project Details
Updated 07/14/26 · Edited by orgFrontier labs, such as Anthropic and Google DeepMind, use activation probes as safety monitors in real world deployed models that users around the world interact with. These probes read the model's internal activations as it processes a request and have become very popular because they are very cheap to train and easy to deploy, and have a high accuracy at detecting whenever a user is trying to generate something harmful in the domains they were trained on.
However, there hasn't been a rigorous study done on how effective these probes are on the rest of the world's languages, especially lower resource ones. Since these probes are predominantly trained and calibrated on English text, and then frozen, if they are failing on other languages, they are failing completely silently (since nothing in the system signals this potential failure, and no one has seriously checked or tested this).
What I will be doing is reimplementing published production probe recipes[1,2] and testing them on a custom dataset that I am currently designing and building. This dataset would have meaning matched harmful and benign phrases across 7 languages around the world, from high (Chinese) to very low resource (Swahili) and be native speaker validated to ensure it is accurate. I will test probes trained from small open source models to frontier open source models, to get as close as possible to the scale of models labs actually deploy.
The final output of this will be a robust, in-depth paper that I will submit to ICLR 2027, as well as to BlackBoxNLP 2026 Workshop Non-Archival Track, a presentation at NEMI(New England Mechanistic Interpretability Workshop) and a released dataset that I hope will be used by safety and interp. researchers around the world, and be built upon beyond the scope of this singular project.
And the ultimate goal of this project is to test whether these monitors fail for lower resource languages, and if they do, show why, and provide a fix that these labs can actually use, so the safety layer works for everyone, not just English speakers.
I am currently doing this project fully solo, with no advisor or lab backing me. I am designing and running everything myself, and paying for compute and native speakers out of my own pocket.
References:
[1] Kramár, J., et al. (2026). Building Production-Ready Probes for Gemini. Google DeepMind. arXiv:2601.11516.
[2] Cunningham, H., et al. (2026). Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks. Anthropic. arXiv:2601.04603.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiThe safety monitors currently in wide use, such as probes, harmfulness classifiers, and SAE-based monitors, are largely trained and evaluated in English, but are then deployed to a global audience, many of whom don't use English. And immediately, the risk is clear: there is a huge hole in oversight that nobody has actually measured.
The safety layer in deployed production models is used by and affects, in one way or another, billions of people, and there has been no serious research done, either by the frontier labs or otherwise, on documenting whether or not the layer actually holds in non English inputs(in particular lower resource and underrepresented languages). If it doesn’t, it could, as we speak, be allowing harmful requests to quietly be let through, without anyone knowing that it’s happening.
And that's what this project aims to explore and document. At the minimum, this project will study the actual gap between how well these monitors work in English vs everything else and if there is a large gap, to do a decomposition of the failures and thorough analysis for why, and give labs a clean eval that they can run on their own monitors. And at a broader level, the aim is to really spread awareness and get more traction on multilingual coverage being a broader part of the AI Safety ecosystem, not just as a robustness check but as a part of the baseline standard for whether a safety system works, so AI truly is safe for everyone using it.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Funding Details
- Jul 13, 2026
- Sep 13, 2026
- 2 months
- -
- -
- -
- -
- -
- -
- -
Private comment. Only shown to approved funders and grant reviewers.