Prior art
What others have done
Before spending a day recording, it pays to map the terrain. Who collected data before us? Who got funded, published, or hired because of it? Which regional programs exist, how much did they cost, and what did they actually ship? This page answers those questions with numbers. The goal is not to be impressed or discouraged by other people's work. The goal is to find the gaps they left.
Who got what
Rewards others have received
The pattern across all of these teams is the same: a clear gap, a specific method, and a first-ever claim. Small teams with community access can compete directly.
| Team | What they did | What they got |
|---|---|---|
| NaijaVoices team | Collected 1,838 hours of Igbo, Hausa, and Yoruba speech | Lacuna Fund grant; Interspeech 2025 paper; a licensable dataset with paid commercial tiers (naijavoices.com) |
| Data Science Nigeria | Ran African Voices, collecting 2,500-3,000 hours of speech across African languages | A share of a $2.2M Gates Foundation grant; featured in Nature |
| Awarri (startup) | Partnered with the Nigerian government to build a national LLM (N-ATLAS), launched at UNGA80 in September 2025 | Government contract; global visibility at a UN launch |
| Intron Health | Built accented clinical speech AI tuned for African-accented English | $1.6M funding; 30+ hospital customers |
| Lelapa AI | African-language AI products and APIs (Vulavula) | $2.5M funding; TIME 100 AI recognition |
| Kallaama (Senegal) | 125 hours of farm-related speech in three Senegalese languages — comparable team size to us | Lacuna Fund grant; published dataset and paper |
| Individual Masakhane contributors | Contributed data and translations to community datasets; no special equipment, no lab affiliation | Co-authorship on top NLP papers that led to scholarships, jobs, and research careers |
| ELDP grantees worldwide | Documented one endangered language per project, any nationality eligible | Up to 10,000 euros each; 500+ projects funded |
Sources: NaijaVoices paper, Gates/DSN grant, Intron TechCrunch, Kallaama arXiv, ELDP grants page
Global programs
What other regions built
Every region that succeeded in language AI had at least one permanent funded institution. Iceland spent roughly $38 per speaker on basic language infrastructure. The Lacuna Fund's total African NLP investment works out to fractions of a cent per speaker. That gap is the context for everything below.
| Program | Money (order of magnitude) | Flagship data | Flagship model | Permanent institution? | Source |
|---|---|---|---|---|---|
| Europe / CLARIN | 1.8M euros (ELE project) plus decades of programme funding | CLARIN repositories across 20+ member countries | National LLMs (varies by country) | Yes (CLARIN, est. 2012) | EU policy |
| Iceland | ~$14M+ for 370,000 speakers (~$38/speaker) | Icelandic Gigaword corpus and speech data | GPT-4 partnership (negotiated directly with OpenAI) | Yes (state programme) | Iceland Review, 2024-26 programme |
| India / Bhashini | ~$10M+ (ministry mission + Nilekani philanthropy) | IndicVoices (7,348 hours); Samanantar (49.7M parallel sentence pairs) | IndicTrans2 covering all 22 official languages | Yes (ministry mission) | Bhashini data report |
| Gulf / ALLaM | $100B-scale pledges (UAE + Saudi Arabia combined, 2024) | 500-billion-token Arabic corpus, built by mobilizing 16 government entities | ALLaM; Jais; Falcon-Arabic | Yes (SDAIA, TII) | Falcon-Arabic |
| SE Asia / SEA-LION | State programme (AI Singapore as anchor) | SEALD corpus, ~1 trillion tokens with SEA languages over-represented | SEA-LION covering 11 languages | Yes (AI Singapore) | SEA-LION paper |
| Latin America / Latam-GPT | $3.5M; 15 countries, 60+ organizations | Regional Spanish/Portuguese corpus | Latam-GPT, launched February 2026 | Forming (Chile CENIA) | Chile gov |
| Africa / Masakhane | ~$10M-scale actual (Lacuna + Gates + Google) versus $60B pledged in the Kigali declaration | ANV ~18,000 hours + WAXAL ~11,000 hours + NaijaVoices 1,838 hours | InkubaLM (0.4B-class); N-ATLAS (8B fine-tune) | No continent-wide body. SADiLaR (South Africa) is the only national example | Lacuna Fund |
Nigeria vs the continent
Where Nigeria sits within Africa
Nigeria leads Africa in language data volume and research output. It trails South Africa in institutions and Kenya and Egypt in government readiness. The pattern is the same as the global one: strong community effort, weak permanent structure.
| Metric | Nigeria | Best African peer | Source |
|---|---|---|---|
| Gov AI Readiness 2025 | 70th globally | Kenya 65th, South Africa 67th | Oxford Insights 2025 |
| Data-centre share of Africa | ~15% (second on the continent) | South Africa ~70% | Introl analysis |
| Open speech data (national languages) | ~4,500+ hours (NaijaVoices + African Voices). Largest on the continent | Rwanda ~2,384 hours (Kinyarwanda via Common Voice) | NaijaVoices paper, Mozilla |
| National LLM | First to launch on the world stage: N-ATLAS at UNGA80, September 2025 | South Africa: none national. InkubaLM is private-sector | ThisDay Live |
| National language-resource institution | None active in AI (NINLAN is silent on the topic) | South Africa's SADiLaR, the continent's only true national language-data infrastructure | SADiLaR |
| Research output | Nigerian authors dominate AfricaNLP contribution rankings | South Africa (DSFSI), Kenya, Ethiopia also strong | Rise of AfricaNLP |
| AI startup capital | Among four countries capturing 83% of Africa's early-2025 AI investment | Kenya, South Africa, Egypt (the other three in the same group) | AI Reports Africa |
| Policy direction | Mother-tongue schooling policy reversed in late 2025, removing a government demand signal for local-language materials | Rwanda and Kenya used government partnerships to create data supply | Global Voices |
Don't undersell this
Where Africa actually leads
Participatory research methodology
Masakhane's model, in which contributors with no lab affiliation co-author papers by contributing data and translations, is now cited and copied globally. No other region invented this at scale. African community researchers proved you don't need a GPU cluster to be a first author.
Kinyarwanda ranked second worldwide on Common Voice
Rwanda's Kinyarwanda community collected roughly 2,384 hours of read speech, placing it second on all of Common Voice worldwide, ahead of French, German, and Spanish communities. A coordinated volunteer effort in a country of 14 million beat the French internet.
30,000+ hours collected in 24 months
Between ANV (~18,000 hours), WAXAL (~11,000 hours), and NaijaVoices (1,838 hours), Africa added more than 30,000 open speech hours in roughly 24 months from 2024 to 2026. No other low-resource region matched this pace. Africa's speech-data position is now surprisingly competitive with India's IndicVoices (7,348 hours).
Code-switching treated as the norm
African datasets like NaijaSenti and AfroCS-xs capture the way people actually speak: mixing languages mid-sentence. Gulf models like Jais only recently started prioritizing this. African NLP researchers modeled real-world language use years before it became a benchmark priority elsewhere.
The landscape shows two things clearly. First, other regions succeeded when they had a permanent institution, a deployment mission, or a sovereign capital base. Africa has the data velocity but not yet the institutions. Second, the gaps left by larger programs are real and specific: only 2.4% of new African speech data comes from speakers aged 50+, the Ebonyi language cluster has zero datasets despite millions of speakers, and NLP work covers fewer than 1% of Nigeria's ~520 languages. Those gaps are where the work is.