Data and downloads

What we publish, and how

Every dataset we release is documented, licensed, and attributed to the communities that made it. This page lists what is out, what is coming, and the terms that govern use.

Seed mark: seed splitting open with a small sprout emerging

License posture

We license so that others can build, but with explicit protections for communities whose voices we recorded.

License posture

  • Text datasets: CC BY-SA 4.0. You may use, adapt, and redistribute with attribution and the same license.
  • Audio datasets: CC BY-SA 4.0 with an additional prohibition on voice cloning and text-to-speech synthesis. This follows the precedent set by the Swivuriso ZA-ANV dataset, which explicitly prohibits TTS and voice-cloning uses.
  • Community ownership: We structure contributor agreements around community ownership and ethical commercialization, consistent with the Esethu framework applied to the ViXSD ASR dataset.
  • BAYIBURU scripture texts: The Izzi, Ezaa, and Ikwo New Testaments (completed 1980) belong to Wycliffe Bible Translators. We do not distribute these texts or readings of them without written permission from Wycliffe. Any dataset derived from scripture readings requires that clearance first.

Datasets

Datasets in progress

NameLanguageKindSizeStatusRelease targetLicense
Ikwo conversational parallel corpusIkwoSpeech + text (parallel)1,000 sentencesIn review2025 Q3CC BY-SA 4.0
Izzi elder-speech pilotIzziSpeech (spontaneous)TBDPlanned2026CC BY-SA 4.0
Ezaa pilotEzaaSpeech (spontaneous)TBDPlanned2026CC BY-SA 4.0
Korring scoping studyKorringScopingTBDPlannedTBDTBD
Archive mark: stacked talking drums representing oral tradition

Where to find our data

HuggingFace

Our primary distribution channel for ML-ready datasets. The org is being set up and releases will appear there first.

olubridge on HuggingFace (planned)

Zenodo

Citable, DOI-backed archival copies for academic use. Each dataset released here gets a stable identifier.

Zenodo community (planned)

ELAR

The Endangered Languages Archive at SOAS takes full deposits, including restricted items that should not appear in open ML datasets. Consent-restricted material goes here so it is preserved without being exposed.

Lanfrica

Lanfrica indexes African-language NLP resources by language and task. Our datasets will be listed there to reach researchers already looking for Ebonyi-region languages.

How to cite

When datasets carry a DOI, use that in your BibTeX entry. Until then, cite the dataset name, version, and this URL.

@misc{olubridge2025ikwo,
  title        = {Ikwo Conversational Parallel Corpus},
  author       = {{Olubridge}},
  year         = {2025},
  note         = {Version 1.0. License: CC BY-SA 4.0 with voice-cloning prohibition},
  howpublished = {\url{https://olubridge.org/data}},
}

A dataset-specific BibTeX entry with DOI will replace this placeholder when the first dataset is released.

Partnering on data

If you are a researcher, a language community group, a university department, or a company that wants to co-collect, co-fund, or build on Ebonyi-language data, we want to hear from you. Tell us your language, your use case, and what you can bring to the table.

We are especially interested in partnerships that put contributors at or above local prevailing rates, consistent with the pay standard set by the Thiomi dataset. If your project pays below that bar, we will talk about closing the gap before we agree.

Send a partnership enquiry