The argument
Why this matters
517 languages, 4 with real NLP
Nigeria has between 517 and 525 living languages depending on how you count dialect continua. Of those, a 2025 survey of Nigerian NLP found that fewer than four have received meaningful research attention. Four. That is under one percent of the country's linguistic reality reflected in the tools its people are increasingly expected to use.
The remaining 513-plus languages are not obscure. Izzi, Ezaa, Ikwo, and Mgbo are spoken across Ebonyi State by communities of hundreds of thousands. They are distinct languages, classified separately by Ethnologue, each with its own ISO code. Izzi and Ezaa are not dialects of Central Igbo. A speaker of standard Igbo from Owerri does not follow a conversation in Izzi without effort. The gap matters because NLP tools built for Igbo do not carry over to these languages, and building tools for Igbo while calling it "Ebonyi language support" is precisely the kind of assumption that makes AI fail the people it is supposed to serve.
The demand is documented
CSA Research surveyed 8,709 consumers across 29 countries and found that 76 percent prefer to buy products when information is available in their own language, and 40 percent will not buy at all from websites in a different language. That is not a preference for linguistic nostalgia. It is a statement about commerce, health information, legal notice, and every other domain where language determines whether you understand the thing in front of you.
Only about 27 percent of sub-Saharan Africans speak English with enough fluency to navigate anglophone digital services. The rest are making do. A survey of platform interface languages found that over 90 percent of Africans must switch to a second language to use major apps. Every time a farmer in Onueke checks a price, every time a patient in Abakaliki reads a discharge summary, that switch is happening.
The demand evidence supports localized services more directly than it supports research artifacts in the abstract. Whoever connects the new datasets to consumer-grade services captures this demand first.
The models fail
The current generation of large language models does not close the gap. The IrokoBench evaluation tested GPT-4o across African languages and found an average score of 59 out of 100. Not a ceiling to push against. A floor that most practical applications cannot survive on. A separate study at WMT 2023 found that ChatGPT loses to older, purpose-built translation systems on 84 percent of low-resource languages. The models are not failing because the problem is too hard. They are failing because the training data does not exist.
AndrĂ¡s Kornai documented this dynamic in 2013: fewer than five percent of the world's languages can still ascend to digital use. His sharpest point for anyone building data today is that archives alone do not create digital life. "Online audio files of an elder reciting folk poetry will not facilitate digital ascent." Communication use by living speakers is the criterion. We are building toward that, not away from it.
Data fixes it
The fix is not theoretical. The NaijaVoices team collected 1,838 hours of speech across Igbo, Hausa, and Yoruba. When they fine-tuned Whisper on that data, word-error rates fell by 76 percent. Not a small improvement. The kind of improvement that makes a product usable where it was not. A doctor's dictation system that was wrong one in four words becomes one that works. A voice search that returned nothing starts returning something.
The same approach works across the scale of effort a small team can bring. The Kallaama project in Senegal collected 125 hours of farm-related speech in three languages on a grant, published a dataset, published a paper, and created a reusable asset that funders can point to. That is the model we are following. The Izzi cluster has no equivalent starting point. The gap is real, the tooling to close it exists, and the first team to do it owns the asset.
What a community's voice is worth
"The demand evidence supports localized services more directly than it supports 'language tech' in the abstract. Whoever connects the new datasets to consumer-grade services captures this demand first."
From the Africa Gaps research, answering whether demand for African-language technology is real and measurable.
Language technology is not a neutral good that distributes itself evenly once invented. It concentrates where the data already exists, which means it concentrates where historical investment already happened. The Izzi cluster enters this era with no text corpus, no speech dataset, no ASR model, no benchmark, and no published paper. That combination, zero coverage across millions of speakers, is exactly what funders reward when a team shows up with a credible plan to fix it. The Ibom NLP project for Efik, Ibibio, Anaang, and Oro became an IJCNLP 2025 paper precisely because the team arrived first.
Elder speech sharpens the case further. A Lanfrica analysis of the African Next Voices datasets found that only 2.4 percent of new African speech data comes from speakers aged 50 and over. That age group is the one that holds idiomatic fluency, oral narrative traditions, and the pronunciation baselines that younger speakers have been code-switching away from for decades. A voice AI that cannot understand a 65-year-old Izzi elder is not a tool for that community. It is a tool for whoever trained it.
Ebonyi specifically
Ebonyi State is usually described in Nigerian discourse as an Igbo-speaking state, and that description is accurate at the level of federal administration. At the level of linguistics it is not. Izzi and Ezaa share around 95 percent of vocabulary with each other and with Ikwo and Mgbo, forming a cluster that linguists treat as a family of related languages. Central Igbo sits outside that cluster. The intelligibility gap is real enough that Ethnologue assigns separate ISO codes and separate Wikipedia articles to each.
Igbo API's dictionary covers 17 dialect variations and includes Izzi and Ezaa entries as pronunciation variants. That is a beginning, not coverage. There is no running text corpus for any of these languages, no speech data, nothing a researcher or engineer could use to train or evaluate a model. A project that produces even a small, clean, well-documented dataset fills a gap that is currently absolute.
The per-speaker funding gap
What governments spend per speaker
Per-speaker calculation methodology: Iceland from Iceland Review. India estimate from Lacuna Fund and public Bhashini programme figures. Africa estimate from Lacuna Fund grant disclosures.
The gap is not a funding problem that will solve itself as AI becomes more widespread. Left unaddressed, the concentration deepens. Models trained on existing data produce outputs that reward the languages already represented, which draws more usage, more feedback data, and more investment to those same languages. Kornai called this digital language death. It happens quietly, and it happens to living languages with living speakers.
The Izzi cluster has millions of speakers, zero NLP resources, a documented funder appetite for exactly this kind of first-dataset work, and a community whose elders are still here. What we are building is described at /plan.