grantmaking.ai Launch Round
We are requesting an ideal grant of $50,000, with a minimum viable funding level of $25,000. The funding would support a focused public-interest research project evaluating whether advanced AI systems maintain the same safety and control standards across African languages that they demonstrate in English.
The funding would not be used for general marketing, consumer app growth, or unrelated commercial product development. It would be used specifically for benchmark development, multilingual safety testing, native-language validation, model evaluation, technical infrastructure, independent review, and publication of the findings.
At the minimum funding level of $25,000, we would build the first version of the benchmark across Yoruba, Swahili, and Nigerian Pidgin. The largest portion, approximately $9,000, would support evaluation pipeline engineering. This includes building the technical system needed to connect to frontier and open-weight AI models, submit multilingual test prompts, store and organize responses, apply scoring criteria, compare behavior across languages, and export the results for analysis. Approximately $4,000 would support research design and benchmark construction, including the threat model, safety categories, prompt templates, scoring rubrics, and methodology. Another $4,500 would support native-language experts who would create, translate, back-translate, culturally adapt, and validate prompts and model outputs.
Approximately $2,500 would be used for frontier-model API access and inference costs, including repeated runs, prompt variations, and comparisons across three to four models. Around $1,500 would support an external AI-safety advisor who would review the theory of impact, threat model, evaluation design, and interpretation of the results. About $1,000 would support secure storage, hosting, and basic benchmark infrastructure. Approximately $500 would support documentation and publication of the technical report, and $2,000 would support project coordination, contractor management, financial administration, and required grant reporting.
With the minimum amount, we expect to produce approximately 600 to 900 multilingual prompts and scenarios. The benchmark would focus on harmful-request refusal, multilingual jailbreaks, instruction following, and translation-based attacks. It would include human validation by native speakers, a reproducible evaluation pipeline, a public benchmark subset, a smaller held-out test set, a technical report, and one virtual briefing for researchers or stakeholders.
At the ideal funding level of $50,000, we would expand the benchmark to five languages by adding Hausa and Igbo. Approximately $14,000 would support more extensive engineering work, including stronger automation, experiment tracking, model-version comparisons, reproducibility controls, and a lightweight public results dashboard. Approximately $6,000 would support expanded research design, allowing us to test additional categories such as deception, oversight failures, code-switching attacks, tool-use and agent-control scenarios, and cases where dangerous capabilities may be accessed differently across languages.
Approximately $10,000 would support native-language prompt creation and validation across all five languages. This would allow for independent second-pass review, adjudication when validators disagree, and basic inter-rater reliability checks. Around $6,000 would cover model API and inference costs, allowing us to test approximately six to eight frontier and open-weight models with more repetitions, adversarial variations, and model-version comparisons.
Approximately $4,000 would support a more involved external AI-safety advisor throughout the design, testing, and reporting stages. Another $3,000 would support dedicated red-team testing and quality assurance, including checks for duplicated prompts, data leakage, inconsistent scoring, weak translations, benchmark contamination, and cases where models appear safe while still providing harmful or misleading assistance.
Approximately $2,500 would support secure storage, hosting, a public benchmark interface, and management of a stronger held-out test set. Another $2,000 would support the technical report, documentation, public release materials, a workshop or webinar, and mitigation guidance for AI developers and policymakers. Approximately $2,500 would support project management, grant administration, contractor coordination, and public quarterly updates.
With the ideal amount, we expect to produce approximately 1,200 to 2,000 prompts and scenarios across five African languages and evaluate six to eight models. The final outputs would include a broader public benchmark, a robust held-out evaluation set, a public model leaderboard or dashboard, a detailed technical report, mitigation recommendations, and a public workshop or webinar.
We would also use NKENNEAi’s partnership with Nigeria’s National Information Technology Development Agency, NITDA, as a pathway to share relevant findings with policymakers and technical stakeholders involved in AI deployment, data governance, and digital infrastructure.
NKENNEAi already has technical infrastructure, language resources, native-speaker relationships, and NSF-supported research experience. This means the grant would not be used to start from zero. It would fund the additional work needed to convert those existing capabilities into a dedicated AI safety evaluation project.
The minimum amount would support a credible first phase that demonstrates whether meaningful multilingual safety gaps exist. The ideal amount would allow us to produce a broader, more rigorous, and more useful public resource for AI laboratories, independent researchers, model deployers, and policymakers.