grantmaking.ai Launch Round
Frontier AI models are deployed to over 200 million Hausa, Yoruba, Igbo, and Nigerian Pidgin speakers across West Africa but they have never been systematically safety-tested in these languages. Anthropic, OpenAI, and Google evaluate their models almost entirely in English. A model that refuses anthrax cultivation instructions or electoral deepfake scripts in English may comply with the same request in Hausa, and no one would know because no one is testing.
This project will builds and publish the first open-source AI safety evaluation benchmark for some African languages, with biosecurity and CBRN as core categories. We will evaluate the latest publicly available frontier models from OpenAI, Anthropic, Google DeepMind, Zhipu AI, DeepSeek, and other leading developers that meet our inclusion criteria at the time of evaluation against 1,125 prompts across 9 harm categories: electoral disinformation, ethnic/religious incitement, health misinformation, financial fraud, violence, privacy, culturally specific harms, biosecurity & dual-use research, and CBRN in Hausa, Yoruba, Igbo, Pidgin, and English (as control). The benchmark will measure the "portability gap": the difference in refusal rates between English and each language.
The project will establish reusable evaluation infrastructure that enables AI developers, researchers, governments, and standards organizations to assess multilingual AI safety consistently across future frontier models in the Global South.
Concrete outputs (all open-source):
- A benchmark dataset of 1,125 prompts (225 per language × 5 languages) authored and reviewed by native speakers, published on Hugging Face
- Model responses from frontier models with harm ratings from two independent native-speaker raters per language
- A technical brief reporting the portability gap findings, with responsible disclosure to model developers (Anthropic, OpenAI, Google) 90 days before public publication, and to Nigerian biosecurity authorities (NBMA, NCDC) immediately for any biosecurity findings
- A published methodology (pre-registered on OSF) so the benchmark can be extended to other African and Global South languages
Who's involved:
The project is led by RAI-GI (Responsible AI Governance Initiative), an independent nonprofit based in Abuja, Nigeria. The 9-person research team — Muhammad Ahmad Janyau (Executive Director, lead), Nigel Hee, Victoria Hyde, Bridget Oviasogie, Bar. Hadiza Makarfi, Shehu Bello Tijjani, Tega Oviasogie, Shamsudden Umar, and Ahmad Ibrahim — includes native speakers of all four target languages. RAI-GI recently published an 75-page baseline assessment, "The State of AI Governance in Nigeria: A Baseline Assessment 2026," which documents the portability gap this project closes (Chapter 14). We are partnered with ForHumanity (independent AI audit and certification) and with Digital Policy Alert, an initiative of the St. Gallen Endowment for Prosperity Through Trade. We are recruiting a biosecurity consultant affiliated with NCDC or NBMA for the biosecurity and CBRN prompt review.
The project can be delivered at two funding levels.
Minimum Funding (US$15,000)
The minimum funding will support the completion of the first public release of the African Multilingual Frontier AI Safety Benchmark. This includes frontier AI model API access and compute, multilingual prompt development across Hausa, Yoruba, Igbo, Nigerian Pidgin, and English, native-language researchers and evaluators, biosecurity and CBRN expert review, statistical analysis and validation, dataset preparation, technical documentation, open-source publication, responsible disclosure to AI developers, project coordination, and contingency for additional API usage or quality assurance requirements.
With this funding, we will deliver an open-source multilingual AI safety benchmark dataset, evaluation results for leading frontier AI models, a technical report, a reproducible evaluation methodology, and responsible disclosure reports for significant safety findings.
Ideal Funding (US$25,000)
The ideal funding includes all activities covered by the minimum budget and expands the project into reusable AI safety infrastructure. Additional funding will support expanding the benchmark to additional African languages, conducting a second evaluation cycle using newer frontier AI models, independent external methodological and statistical review, development of a public benchmark dashboard and leaderboard, additional research assistance and data engineering support, benchmark maintenance and open-source updates, and broader dissemination through workshops and stakeholder engagement with AI developers, policymakers, researchers, and standards organizations.
The requested funding will be allocated approximately as follows:
-
Frontier AI model API access and compute: US$3,500
-
Multilingual prompt development and validation: US$2,500
-
Native-language researchers, evaluators and quality assurance: US$5,500
-
Biosecurity and CBRN expert review: US$2,000
-
Statistical analysis, validation and external methodological review: US$2,500
-
Dataset engineering, documentation, publication and benchmark dashboard: US$2,500
-
Project management, research coordination and administration: US$2,000
-
Responsible disclosure, dissemination and stakeholder engagement: US$1,000
-
Benchmark maintenance, open-source updates and contingency: US$1,000
This funding will establish the first reusable multilingual frontier AI safety benchmark for African languages and provide open evaluation infrastructure that researchers, AI developers, governments, and standards organizations can use to assess and improve the safety of frontier AI systems across multilingual contexts.