Frontier model safety training such as refusals, dangerous-capability guardrails, jailbreak resistance is developed and evaluated overwhelmingly in English. There is growing evidence that these safeguards degrade sharply in low-resource languages, and that translating a harmful request into a low-resource language is itself a working jailbreak vector. As frontier models are deployed to hundreds of millions of Bengali speakers (and across South Asia), this creates a real, largely unmeasured misuse surface.
I will build the first systematic safety-evaluation benchmark for Bengali, covering (1) refusal robustness on harmful requests, (2) low-resource-language jailbreak attacks, and (3) dangerous-capability elicitation compared against English baselines. I'll evaluate leading models (Claude, GPT, Gemini, Llama), quantify the English-vs-Bengali safety gap, and responsibly disclose exploitable failures to the relevant labs before any public release. Outputs: an open-source benchmark and dataset (HuggingFace), reproducible evaluation code, and a public technical report.
I'm a Bangladeshi ML researcher (Brac University) with published transformer-based Bengali NLP work and 360+ citations across deep-learning and evaluation research. As a native Bengali speaker embedded in this research community, I can construct linguistically valid adversarial data that non-speakers cannot — the core reason this gap remains unmeasured. If successful, I'll extend to Urdu, Nepali, and Sinhala as a reusable multilingual safety-eval framework.