Prior art

What others have done

Before spending a day recording, it pays to map the terrain. Who collected data before us? Who got funded, published, or hired because of it? Which regional programs exist, how much did they cost, and what did they actually ship? This page answers those questions with numbers. The goal is not to be impressed or discouraged by other people's work. The goal is to find the gaps they left.

Who got what

Rewards others have received

The pattern across all of these teams is the same: a clear gap, a specific method, and a first-ever claim. Small teams with community access can compete directly.

TeamWhat they didWhat they got
NaijaVoices teamCollected 1,838 hours of Igbo, Hausa, and Yoruba speechLacuna Fund grant; Interspeech 2025 paper; a licensable dataset with paid commercial tiers (naijavoices.com)
Data Science NigeriaRan African Voices, collecting 2,500-3,000 hours of speech across African languagesA share of a $2.2M Gates Foundation grant; featured in Nature
Awarri (startup)Partnered with the Nigerian government to build a national LLM (N-ATLAS), launched at UNGA80 in September 2025Government contract; global visibility at a UN launch
Intron HealthBuilt accented clinical speech AI tuned for African-accented English$1.6M funding; 30+ hospital customers
Lelapa AIAfrican-language AI products and APIs (Vulavula)$2.5M funding; TIME 100 AI recognition
Kallaama (Senegal)125 hours of farm-related speech in three Senegalese languages — comparable team size to usLacuna Fund grant; published dataset and paper
Individual Masakhane contributorsContributed data and translations to community datasets; no special equipment, no lab affiliationCo-authorship on top NLP papers that led to scholarships, jobs, and research careers
ELDP grantees worldwideDocumented one endangered language per project, any nationality eligibleUp to 10,000 euros each; 500+ projects funded

Sources: NaijaVoices paper, Gates/DSN grant, Intron TechCrunch, Kallaama arXiv, ELDP grants page

Global programs

What other regions built

Every region that succeeded in language AI had at least one permanent funded institution. Iceland spent roughly $38 per speaker on basic language infrastructure. The Lacuna Fund's total African NLP investment works out to fractions of a cent per speaker. That gap is the context for everything below.

ProgramMoney (order of magnitude)Flagship dataFlagship modelPermanent institution?Source
Europe / CLARIN1.8M euros (ELE project) plus decades of programme fundingCLARIN repositories across 20+ member countriesNational LLMs (varies by country)Yes (CLARIN, est. 2012)EU policy
Iceland~$14M+ for 370,000 speakers (~$38/speaker)Icelandic Gigaword corpus and speech dataGPT-4 partnership (negotiated directly with OpenAI)Yes (state programme)Iceland Review, 2024-26 programme
India / Bhashini~$10M+ (ministry mission + Nilekani philanthropy)IndicVoices (7,348 hours); Samanantar (49.7M parallel sentence pairs)IndicTrans2 covering all 22 official languagesYes (ministry mission)Bhashini data report
Gulf / ALLaM$100B-scale pledges (UAE + Saudi Arabia combined, 2024)500-billion-token Arabic corpus, built by mobilizing 16 government entitiesALLaM; Jais; Falcon-ArabicYes (SDAIA, TII)Falcon-Arabic
SE Asia / SEA-LIONState programme (AI Singapore as anchor)SEALD corpus, ~1 trillion tokens with SEA languages over-representedSEA-LION covering 11 languagesYes (AI Singapore)SEA-LION paper
Latin America / Latam-GPT$3.5M; 15 countries, 60+ organizationsRegional Spanish/Portuguese corpusLatam-GPT, launched February 2026Forming (Chile CENIA)Chile gov
Africa / Masakhane~$10M-scale actual (Lacuna + Gates + Google) versus $60B pledged in the Kigali declarationANV ~18,000 hours + WAXAL ~11,000 hours + NaijaVoices 1,838 hoursInkubaLM (0.4B-class); N-ATLAS (8B fine-tune)No continent-wide body. SADiLaR (South Africa) is the only national exampleLacuna Fund

Nigeria vs the continent

Where Nigeria sits within Africa

Nigeria leads Africa in language data volume and research output. It trails South Africa in institutions and Kenya and Egypt in government readiness. The pattern is the same as the global one: strong community effort, weak permanent structure.

MetricNigeriaBest African peerSource
Gov AI Readiness 202570th globallyKenya 65th, South Africa 67thOxford Insights 2025
Data-centre share of Africa~15% (second on the continent)South Africa ~70%Introl analysis
Open speech data (national languages)~4,500+ hours (NaijaVoices + African Voices). Largest on the continentRwanda ~2,384 hours (Kinyarwanda via Common Voice)NaijaVoices paper, Mozilla
National LLMFirst to launch on the world stage: N-ATLAS at UNGA80, September 2025South Africa: none national. InkubaLM is private-sectorThisDay Live
National language-resource institutionNone active in AI (NINLAN is silent on the topic)South Africa's SADiLaR, the continent's only true national language-data infrastructureSADiLaR
Research outputNigerian authors dominate AfricaNLP contribution rankingsSouth Africa (DSFSI), Kenya, Ethiopia also strongRise of AfricaNLP
AI startup capitalAmong four countries capturing 83% of Africa's early-2025 AI investmentKenya, South Africa, Egypt (the other three in the same group)AI Reports Africa
Policy directionMother-tongue schooling policy reversed in late 2025, removing a government demand signal for local-language materialsRwanda and Kenya used government partnerships to create data supplyGlobal Voices

Don't undersell this

Where Africa actually leads

Bridge mark: arch bridge, the Olubridge wordmark accent

Participatory research methodology

Masakhane's model, in which contributors with no lab affiliation co-author papers by contributing data and translations, is now cited and copied globally. No other region invented this at scale. African community researchers proved you don't need a GPU cluster to be a first author.

Bridge mark: arch bridge, the Olubridge wordmark accent

Kinyarwanda ranked second worldwide on Common Voice

Rwanda's Kinyarwanda community collected roughly 2,384 hours of read speech, placing it second on all of Common Voice worldwide, ahead of French, German, and Spanish communities. A coordinated volunteer effort in a country of 14 million beat the French internet.

Bridge mark: arch bridge, the Olubridge wordmark accent

30,000+ hours collected in 24 months

Between ANV (~18,000 hours), WAXAL (~11,000 hours), and NaijaVoices (1,838 hours), Africa added more than 30,000 open speech hours in roughly 24 months from 2024 to 2026. No other low-resource region matched this pace. Africa's speech-data position is now surprisingly competitive with India's IndicVoices (7,348 hours).

Bridge mark: arch bridge, the Olubridge wordmark accent

Code-switching treated as the norm

African datasets like NaijaSenti and AfroCS-xs capture the way people actually speak: mixing languages mid-sentence. Gulf models like Jais only recently started prioritizing this. African NLP researchers modeled real-world language use years before it became a benchmark priority elsewhere.

The landscape shows two things clearly. First, other regions succeeded when they had a permanent institution, a deployment mission, or a sovereign capital base. Africa has the data velocity but not yet the institutions. Second, the gaps left by larger programs are real and specific: only 2.4% of new African speech data comes from speakers aged 50+, the Ebonyi language cluster has zero datasets despite millions of speakers, and NLP work covers fewer than 1% of Nigeria's ~520 languages. Those gaps are where the work is.