The surprising finding

Ebonyi speaks languages the world has never recorded

Most people outside Ebonyi State assume the population speaks Igbo. Linguists disagree. Izzi, Ezaa, Ikwo, and Mgbo form a cluster with roughly 95% shared vocabulary among themselves, but only marginal intelligibility with Central Igbo. Ethnologue treats each one as a distinct language, with its own Wikipedia article and ISO 639-3 code. The cluster has millions of speakers. It has no text corpus, no automatic speech recognition system, and no text-to-speech model anywhere on earth.

Ebonyi also hosts Korring (spoken by the Orring people), which is not Igboid at all. It belongs to the Cross-River / semi-Bantoid family, making it the most linguistically isolated language in the state and, arguably, the most scientifically valuable. Beyond the main cluster, Afikpo (Ehugbo), Edda, Okposi, and Uburu each carry distinct speech varieties. None has dedicated AI resources. Together, they represent the sharpest gap we have found in all of Nigeria's 517 or so documented languages: millions of speakers, zero coverage, and a community that has never once been asked.

The cluster

Five languages, zero AI resources

Izzi (Izi)

izz

~540,000 speakers

What is missing

  • Text corpus
  • Automatic speech recognition
  • Text-to-speech system
  • Any benchmarked NLP model

Ezaa

eza

What is missing

  • Text corpus
  • Automatic speech recognition
  • Text-to-speech system
  • Any benchmarked NLP model

Ikwo

iqw

What is missing

  • Open text corpus
  • Automatic speech recognition
  • Text-to-speech system
  • Any benchmarked NLP model

A draft 1,000-sentence conversational parallel corpus is in review by our team.

Mgbo

gmz

What is missing

  • Any Bible text (audio exists; text-side gap)
  • Text corpus
  • Automatic speech recognition
  • Text-to-speech system

Korring / Orring

krh

What is missing

  • Scripture text or audio (none found)
  • Any dictionary or word list
  • Text corpus
  • Automatic speech recognition
  • Text-to-speech system

Not Igboid at all. A Cross-River / semi-Bantoid language. The deepest gap in the state.

Also in Ebonyi

Further varieties

These communities likely use the standard Igbo Bible. Their local speech varieties have no dedicated AI resources and have not been systematically documented.

Afikpo (Ehugbo)

Spoken in Afikpo North and South LGAs. No dedicated scripture or NLP resource found.

Edda

Spoken in Afikpo South LGA. Pronunciation variants catalogued in Nkowa okwu only.

Okposi

Spoken in Ohaozara LGA. No dedicated resource found.

Uburu

Spoken in Ohaozara LGA alongside Okposi. No dedicated resource found.

Existing anchors

What the community already built

Archive mark: stacked talking drums representing oral tradition

Scripture texts

The Ezaa, Izii, and Ikwo communities received their New Testaments in 1980. All three are free on YouVersion today. The rights belong to Wycliffe Bible Translators; any corpus built from people reading these texts needs written permission first.

Voice mark: sound waves from an open mouth

Gospel audio

Global Recordings Network holds free MP3 recordings in Izii, Ezaa, Ikwo, and Mgbo. That is four of the five main cluster languages. Korring has nothing.

Elder mark: dignified person in headscarf silhouette

Dialect word lists

The Igbo API dialect docs include Izzi, Ezaa, and Edda as pronunciation variants. These are word-level entries only, not sentences or speech. They are a seed lexicon, not a corpus.

Why this matters for AI

First dataset, guaranteed paper

The pattern is documented. When Ibom NLP created the first dataset for Efik, Ibibio, Anaang, and Oro, it became an IJCNLP 2025 paper precisely because nobody had done it before. The paper's leverage came from the zero-baseline gap, not from algorithmic novelty.

The Izzi cluster is in the same position today. Millions of speakers. No ASR, no TTS, no benchmark. The combination is rare: high speaker count, total absence of resources, and a research community that actively rewards first-mover datasets. GPT-4o scores roughly 59 out of 100 on African-language benchmarks. Whisper's error rate drops by 76% when fine-tuned on 1,838 hours of NaijaVoices speech. For a language with zero current data, any hours collected are publishable. A 125-hour first corpus, like Kallaama in Senegal, is enough.

The documentation funders know this. ELDP pays up to 10,000 euros per grant for a first-ever language documentation. Lacuna Fund's agriculture and health rounds list under-resourced African languages as a standing priority. The verified gap in elder speech data (only 2.4% of new African speech data comes from speakers over 50) makes an oral-history-style Ebonyi collection fundable on two separate grounds at once.

“Onye wetara oji wetara ndu”

Whoever brings kola nut brings life.

Izi proverb, Ebonyi State