Measuring AI Intelligence in African Contexts.

The Datalens Africa LLM Leaderboard is the definitive benchmark for evaluating AI understanding of African contexts. Compare how leading AI systems perform across language, culture, knowledge, and healthcare using the most comprehensive African evaluation framework.

15k+

Questions

16+

Languages

32

Specialties

5

LLM Families

African AI Benchmark Leaderboard

← Swipe to see all benchmark scores →
Rank
Model
AfriMCQA
AfriMMLU
MasakhaNEWS
AfriMedQA
Overall Score
👑1
Gemini Logo
Gemini 3.7 Flash
Google · Aug '26
91.11%
82.75%
77.06%
82.65%
83.39%
2
Gemini Logo
Gemini 3.5 Flash
Google · May '26
90.14%
80.88%
72.74%
84.71%
82.12%
3
OpenAI Logo
GPT-5.6 Sol
OpenAI · Aug '26
84.38%
78.13%
82.88%
75.93%
80.33%
4
Grok Logo
Grok 4.6
xAI · Sep '26
83.65%
77.50%
74.61%
73.31%
77.27%
5
Claude Logo
Claude Opus 4.6
Anthropic · May '26
82.45%
68.13%
79.19%
78.99%
77.19%
6
DeepSeek Logo
DeepSeek-V4-Pro
DeepSeek AI · May '26
76.75%
75.80%
76.24%
76.26%
7
OpenAI Logo
GPT-5.4
OpenAI · May '26
78.37%
72.75%
77.71%
74.71%
75.88%
8
DeepSeek Logo
DeepSeek-V4-Flash
DeepSeek AI · Aug '26
77.75%
79.87%
69.26%
75.63%
9
OpenAI Logo
GPT-5.1
OpenAI · Feb '26
81.22%
62.13%
78.61%
76.72%
74.67%
10
Gemini Logo
Gemini 3.1 Flash Lite
Google · May '26
85.58%
69.38%
71.44%
72.06%
74.61%
11
Claude Logo
Claude Sonnet 5
Anthropic · Aug '26
75.72%
55.50%
82.62%
78.76%
73.15%
12
Gemini Logo
Gemini 2.5 Pro
Google · Feb '26
85.58%
79.47%
51.22%
75.95%
73.05%
13
OpenAI Logo
GPT-5.6 Luna
OpenAI · Aug '26
68.51%
65.37%
81.25%
76.40%
72.88%
14
DeepSeek Logo
DeepSeek-R1
DeepSeek AI · Feb '26
70.63%
73.87%
72.99%
72.50%
15
Claude Logo
Claude Sonnet 4.6
Anthropic · Feb '26
79.33%
62.38%
68.75%
78.04%
72.12%
16
OpenAI Logo
GPT-5.2
OpenAI · Feb '26
75.50%
67.38%
75.27%
69.97%
72.03%
17
Gemini Logo
Gemini 2.5 Flash
Google · Feb '26
81.05%
60.50%
71.20%
74.42%
71.79%
18
Grok Logo
Grok 4.1 Fast Reasoning
xAI · Feb '26
72.60%
59.63%
64.90%
70.13%
66.81%
19
Claude Logo
Claude Opus 5
Anthropic · Aug '26
59.37%
51.00%
74.45%
79.39%
66.05%
20
DeepSeek Logo
DeepSeek-V3.2
DeepSeek AI · Feb '26
61.50%
64.66%
71.06%
65.74%
21
Grok Logo
Grok 4 Fast Reasoning
xAI · Feb '26
74.76%
54.13%
60.26%
71.32%
65.12%
22
Claude Logo
Claude Haiku 4.5
Anthropic · Feb '26
63.46%
54.75%
67.72%
62.49%
62.10%
+
Your model here
Submit a model and we return a scored, reproducible evaluation across seven dimensions.

* Scores represent average performance across all benchmark sub-tasks in each dataset. Version 1.4 · Last updated: 1 September 2026.

AfriMCQA
Multilingual Cultural Understanding · Accuracy (%)
AfriMMLU
Knowledge & Reasoning Across Languages · Accuracy (%)
MasakhaNEWS
News Topic Classification · Macro-F1 (%)
AfriMedQA v2
Pan-African Clinical QA · Macro-Accuracy (%)

Four Pillars of African AI Evaluation

Each benchmark targets a distinct dimension of African contextual understanding, spanning culture, language, and clinical medicine, and providing a multidimensional view of LLM capability.

01
Afri-MCQA
Multimodal, culturally-grounded multiple-choice questions in 16+ African languages, testing deep cultural and linguistic comprehension.
Culture Multilingual Multimodal
Languages 16+
Task Type MCQ / VQA
Metric Accuracy
02
AfriMMLU
Human-translated MMLU spanning 5 subjects across 17 African languages. Tests knowledge across geography, law, economics, global facts and mathematics.
Knowledge Reasoning Human-Translated
Subjects 5 Domains
Languages 17
Metric Accuracy
03
MasakhaNEWS
News topic classification across 16 African languages and 7 categories. Tests in-language understanding of real-world African news across diverse topics.
NLP Classification News
Categories 7 Topics
Languages 16
Metric Macro-F1
04
AfriMedQA v2
Pan-African medical QA dataset with 15,275 questions from 621 contributors across 16 countries, covering 32+ clinical specialties.
Medical Expert-Curated Healthcare
Questions 15,275
Specialties 32+
Metric Macro-Accuracy

Frontier AI Systems Under Review

Google
Gemini Series
Five Gemini models evaluated; Gemini 3.7 Flash leads the entire leaderboard with the highest overall score.
Models3.7 Flash · 3.5 Flash
Best Score83.39%
OpenAI
GPT Series
Five GPT models evaluated across all four African benchmarks; GPT-5.6 Sol is the strongest, leading on MasakhaNEWS.
Models5.1 · 5.2 · 5.4 · 5.6 Sol · 5.6 Luna
Best Score80.33%
xAI
Grok Series
Three Grok models evaluated with web access disabled for fairness; Grok 4.6 jumps the family to #4 overall, well ahead of the Fast Reasoning models.
ModelsGrok 4 · Grok 4.1 · Grok 4.6
Best Score77.27%
Anthropic
Claude Series
Five Claude models evaluated; Opus 4.6 leads the family and ranks #5 overall, with Sonnet 5 close behind on MasakhaNEWS.
ModelsOpus 4.6/5 · Sonnet 4.6/5 · Haiku 4.5
Best Score77.19%
DeepSeek AI
DeepSeek Series
DeepSeek-V4-Pro, V4-Flash, R1, and V3.2 evaluated; V4-Pro leads the family with strong MMLU and MasakhaNEWS scores.
ModelsV4-Pro · V4-Flash · R1 · V3.2
Best Score76.26%

This leaderboard ranks the models we chose. Get yours measured.

The four benchmarks here are one dimension of seven. Submit your model and we return a scored, reproducible report — per language, with a documented methodology.

Submitting costs nothing and commits you to nothing — we reply with a scoped, costed quote.

1
African-language task performance — this leaderboard
2
Token efficiency / cost-to-serve
3
Cultural appropriateness
4
Factuality & hallucination rate
5
Safety in African contexts
6
Instruction following & reasoning
7
Dialect & code-switch robustness
DataLens Studio

Help the next model
top this leaderboard.

Every score on this leaderboard is a reflection of the data behind the model. Better African language annotations produce better training data, which trains models that score higher on these benchmarks. DataLens Studio is the platform where that work happens — built specifically for African language labeling, RLHF, and cultural context annotation.

1
Annotate African language data
Label text, audio, and cultural context across 50+ African languages.
2
Build better training datasets
Human-verified, culturally-grounded data that frontier models are missing.
3
Models score higher on African benchmarks
Your contributions directly move the needle on this leaderboard.
DataLens
DataLens Studio
African Language Annotation Platform
Text, audio & RLHF annotation tasks
50+ African languages supported
Earn while contributing to African AI
Quality-reviewed by African language experts

Every annotation is a step toward AI that truly understands Africa.

Submit your model for evaluation

We review every submission by hand and reply with a scoped, costed quote. If your model is not yet public, say so in the notes and we will work under NDA.

Pick the languages that matter commercially. Leave blank and we will propose a default set based on your market.
All seven are selected by default. Narrowing the scope lowers the cost.

Submitting costs nothing and commits you to nothing. We reply with a scoped, costed quote — usually within three working days.

Submission received.

An evaluation lead will review your model against the dimensions and languages you selected, and come back with a costed quote and timeline — usually within three working days. If it is urgent, book a slot below and we will scope it live.

Book a scoping call