How well do LLMs understand Nigeria?
Open scores for how well LLMs perform in different downstream tasks in Hausa, Igbo, Yoruba, and Pidgin. i
Overall
iBreakdown
click a column to sorti
vs
Head to head
iShow every language and the statistics ▾
Score difference on the same questions. Bold = statistically clear (paired bootstrap, 95%).
English vs Nigerian languages
iExample answers
Models
What is this?
A repeatable benchmark of language models on Nigerian languages. N-ATLAS, Nigeria's open LLM, is the current spotlight.
How are scores calculated?
- Full test sets, limited to Hausa, Igbo, Yoruba, Pidgin and English. Same prompts for every model.
- Each benchmark keeps its original metric: weighted F1 (NaijaSenti, MasakhaNEWS), accuracy (AfriXNLI, AfriMMLU, MMLU-Pro), exact match (AfriMGSM), chrF++ (FLORES+ v4.7), strict accuracy (IFEval), refusal rate (LSR).
- Task scores on the leaderboard are the plain average of their benchmarks. Overall averages the task categories built on Nigerian and African-language benchmarks; English-only benchmarks (IFEval, MMLU-Pro) are not counted in it. "Nigerian" means the average of Hausa, Igbo and Yoruba.
- Every score has a 95% bootstrap confidence interval; comparisons are paired on the same questions.
How were the models run?
- Open models: locally with MLX at 8-bit precision, greedy decoding, no thinking mode.
- Gemini 3.1 Pro: batch API, provider-default sampling, thinking LOW (it cannot be turned off).
- Answers are parsed by rule. LSR safety answers are graded by GPT (gpt-6-astra), a provider none of the graded models come from.
What should I keep in mind?
- Gemini's NaijaSenti uses a fixed 1,000-tweet sample per language; its MMLU-Pro is the published Vals AI figure (shown in italics, not ranked).
- Requests that failed at the API are excluded rather than scored wrong (0.7% of Gemini requests).
- LSR has 14 prompts per language, so its intervals are wide. The score counts only clear refusals; answers that show the model did not understand the request (gibberish, wrong language) are counted separately, not as refusals.
- Pidgin and Igala are not in N-ATLAS's training data.
- Public benchmarks may appear in any model's training data.