Amharic AI Benchmark
How Well Do Leading AI Models Understand Amharic?
Less well than their marketing suggests, and far less well than they perform in English. For organizations betting on Amharic‑language AI, the headline capability numbers in product announcements do not transfer.
Across the strongest independent benchmarks for African languages, today’s leading models show a wide and consistent gap between high‑resource languages like English and low‑resource languages like Amharic, and the best freely available models trail the best commercial ones by a large margin. What works in English quietly breaks in Amharic, and the only reliable way to know how badly is to test.
This brief summarizes what the public evidence actually shows as of mid‑2026, where that evidence is strong, and where it is still missing.
01Why this question is hard to answer well
Most state‑of‑the‑art models are trained overwhelmingly on high‑resource languages, so they tend to underperform on languages that were thinly represented, or effectively absent, during training. Amharic, despite being spoken by tens of millions of people, sits firmly in the low‑resource category for machine learning: there is comparatively little high‑quality digital text and audio, and the language’s Ge’ez‑based script and rich morphology make it genuinely harder to model than the European languages these systems were optimized around.
There is a second, subtler problem. Amharic‑specific public evaluation is itself scarce. Many low‑resource languages have historically been tested only on simple text‑classification tasks, because comprehensive benchmarks did not exist. That means a lot of confident claims about model performance in Amharic rest on very little measurement. The gap in evaluation is part of the story, and part of why independent, Amharic‑grounded assessment matters.
02What the benchmarks show
The most credible recent measurement comes from IrokoBench, a human‑translated benchmark covering more than a dozen typologically diverse African languages, including Amharic, across three demanding tasks: natural language inference, grade‑school mathematical reasoning, and multiple‑choice knowledge questions (AfriXNLI, AfriMGSM, and AfriMMLU respectively). It was produced by researchers from Masakhane, University College London, Lelapa AI, Cohere For AI, Microsoft Research Africa, and Ethiopia’s own Haramaya University, and published at NAACL in 2025.
Two findings stand out, and both matter for anyone planning to deploy Amharic AI.
- A large high‑resource‑to‑African‑language gap. Models that look strong in English degrade sharply when the same tasks are posed in African languages. The capability you pay for in English is not the capability you get in Amharic.
- A large open‑versus‑proprietary gap. The best‑performing open model in the study reached only about 58 percent of the score of the best proprietary model. In practice, teams reaching for a free, self‑hostable model to handle Amharic are often starting from a much weaker baseline than they realize.
In plain terms
IrokoBench reports results across African languages as a group, and Amharic is one language within it. The aggregate findings are robust and directly relevant, but granular, Amharic‑only public leaderboards remain thin. That measurement gap is itself a finding.
What works in English quietly breaks in Amharic.
03The spoken‑language picture is the same, or worse
Text is only half of real‑world Amharic AI. Speech matters enormously, given Amharic’s strong oral culture and the practical value of voice interfaces where literacy and keyboard input are barriers.
Here the evidence is direct. Research fine‑tuning OpenAI’s Whisper speech‑recognition model for Amharic found that the off‑the‑shelf model struggles with the language because it was thinly represented in training, and that meaningful accuracy requires fine‑tuning on Amharic‑specific data, drawing on resources like Mozilla Common Voice, FLEURS, and locally collected speech corpora. Notably, the same work found that handling an Amharic‑specific linguistic feature, the normalization of homophones, measurably improved transcription accuracy. That is a concrete example of a broader truth: closing the gap requires linguistic knowledge of Amharic, not just more compute.
04What Ethiopia’s own builders are demonstrating
The most instructive evidence comes from teams that have chosen to build for Amharic specifically, rather than wait for general models to catch up.
The clearest case is Lesan AI, founded in 2019, which built dedicated machine translation for Amharic and Tigrinya. Through human evaluation, Lesan’s system has been shown to outperform general‑purpose offerings from Google and Microsoft on Ethiopian language pairs. Their approach is telling: rather than relying only on web‑scraped data, they built from offline print resources and a custom OCR system designed for the Ethiopic script, working with a large network of human translators. Lesan’s co‑founder has put the alternative bluntly, observing that general chatbots are effectively useless for these languages, returning nonsense or words that do not exist.
A wider Ethiopian ecosystem is forming around the same insight that data and linguistic depth come first: efforts such as Leyu (ለዩ), crowdsourcing labeled datasets for Amharic and other Ethiopian languages, and longer‑standing players like iCog Labs, point to a strategy built on language‑specific data and human expertise, rather than the hope that a single global model will solve the problem.
The pattern is consistent, and it is the central finding of this brief. Specialized, linguistically grounded, human‑in‑the‑loop approaches outperform general‑purpose models on Amharic, and they do so precisely because they treat Amharic as a language with its own structure, not as a translation problem to be solved from English.
05What this means if you are building or buying Amharic AI
- Do not trust English‑derived capability claims. A model’s benchmark scores and demos are almost always English‑first. Assume Amharic performance is materially lower until measured.
- Open models are not a safe default for Amharic. The open‑versus‑proprietary gap is large, so “we will just use an open model” can mean starting far behind.
- Evaluation is not optional. Because Amharic‑specific public benchmarks are limited, off‑the‑shelf leaderboards will not tell you what you need to know. Task‑specific, Amharic‑native evaluation is the only reliable signal.
- Human linguistic expertise is the differentiator. The teams getting results combine models with native Amharic knowledge, for data, for evaluation, and for correction. Machines help, but humans lead.
06How Amharic Intelligence approaches this
We evaluate AI systems for Amharic the way the evidence says they should be evaluated: with Amharic‑native linguistic expertise, task‑specific testing rather than borrowed English benchmarks, and human‑in‑the‑loop review that catches the failures automated scores miss. Whether you are training a model, localizing a product, or auditing an AI‑translated corpus, we help you find out how well your system actually understands Amharic, and close the gap where it does not.
Sources
- Adelani et al., IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models, NAACL 2025 (arXiv:2406.03368).
- Whispering in Amharic: Fine-tuning Whisper for Low-resource Language, 2025 (arXiv:2503.18485).
- Lesan AI company materials and founder interviews, with ecosystem reporting on Ethiopian AI (2023 to 2025).
- FLEURS and Mozilla Common Voice, referenced low-resource speech datasets.
This brief synthesizes publicly available research and reporting; figures are drawn from the cited sources. Amharic-specific public benchmarking remains an active gap, and we update our analysis as new evidence emerges.
