grantmaking.ai Launch Round
Frontier labs, such as Anthropic and Google DeepMind, use activation probes as safety monitors in real world deployed models that users around the world interact with. These probes read the model's internal activations as it processes a request and have become very popular because they are very cheap to train and easy to deploy, and have a high accuracy at detecting whenever a user is trying to generate something harmful in the domains they were trained on.
However, there hasn't been a rigorous study done on how effective these probes are on the rest of the world's languages, especially lower resource ones. Since these probes are predominantly trained and calibrated on English text, and then frozen, if they are failing on other languages, they are failing completely silently (since nothing in the system signals this potential failure, and no one has seriously checked or tested this).
What I will be doing is reimplementing published production probe recipes[1,2] and testing them on a custom dataset that I am currently designing and building. This dataset would have meaning matched harmful and benign phrases across 7 languages around the world, from high (Chinese) to very low resource (Swahili) and be native speaker validated to ensure it is accurate. I will test probes trained from small open source models to frontier open source models, to get as close as possible to the scale of models labs actually deploy.
The final output of this will be a robust, in-depth paper that I will submit to ICLR 2027, as well as to BlackBoxNLP 2026 Workshop Non-Archival Track, a presentation at NEMI(New England Mechanistic Interpretability Workshop) and a released dataset that I hope will be used by safety and interp. researchers around the world, and be built upon beyond the scope of this singular project.
And the ultimate goal of this project is to test whether these monitors fail for lower resource languages, and if they do, show why, and provide a fix that these labs can actually use, so the safety layer works for everyone, not just English speakers.
I am currently doing this project fully solo, with no advisor or lab backing me. I am designing and running everything myself, and paying for compute and native speakers out of my own pocket.
References:
[1] Kramár, J., et al. (2026). Building Production-Ready Probes for Gemini. Google DeepMind. arXiv:2601.11516.
[2] Sharma, M., et al. (2025). Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming. Anthropic. arXiv:2501.18837.
Minimum Cost:
- Native Speaker Validation: 12 validators across 6 languages, $30/hr, 400 items each[$4300]
- Compute and API costs (activation extraction, translation pipeline, and GPU rental for frontier-scale open model experiments if free NDIF access is delayed)[$900]
- Contingency(validator replacement, translation re-runs, compute overages)[$500]
Ideal Amount:
- Native Speaker Validation: 12 validators across 6 languages, $30/hr, 400 items each with including a re-validation buffer[$4800]
- Travel to present this work at the NEMI interpretability workshop in Boston (Aug 14) [$700]
- Full Possible Compute Cost, including GPU rental fallback for frontier-scale open model experiments, activation storage across all models, and re-runs if any issues or errors occur [$1400]
- Contingency is the same [$500]
Private comment. Only shown to approved funders and grant reviewers.