Uhuru's flagship, Helios 5.0, measured head-to-head with the strongest external frontier model on African-language benchmarks — identical questions, identical conditions. On African-language mathematical reasoning it comes out ahead across every African language tested, by a statistically significant margin (McNemar p = 0.023) — and a second maths benchmark corroborates it. On reading comprehension the two are level. Every number here is measured and reproducible; re-run it yourself.
| Language | Helios 5.0 Uhuru AI · flagship | Helios 4.0 Uhuru AI · agentic | Claude Opus 4.8 Anthropic |
|---|---|---|---|
| Swahili | 96.0% | 93.3% | 95.3% |
| Amharic | 93.9% | 81.0% | 93.2% |
| Yoruba | 91.1% | 80.8% | 90.4% |
| Hausa | 89.5% | 80.4% | 87.4% |
| Shona | 87.0% | 76.7% | 84.2% |
| isiZulu | 81.5% | 71.2% | 78.1% |
| isiXhosa | 78.7% | 65.2% | 75.9% |
| English (control) | 98.6% | 100.0% | 99.3% |
| Macro average | 89.5% | 81.1% | 88.0% |
AfriMGSM · masakhane/afrimgsm (IrokoBench)· n≈150/language · seed 20260719 · temperature 0