Data and downloads
What we publish, and how
Every dataset we release is documented, licensed, and attributed to the communities that made it. This page lists what is out, what is coming, and the terms that govern use.
License posture
We license so that others can build, but with explicit protections for communities whose voices we recorded.
License posture
- Text datasets: CC BY-SA 4.0. You may use, adapt, and redistribute with attribution and the same license.
- Audio datasets: CC BY-SA 4.0 with an additional prohibition on voice cloning and text-to-speech synthesis. This follows the precedent set by the Swivuriso ZA-ANV dataset, which explicitly prohibits TTS and voice-cloning uses.
- Community ownership: We structure contributor agreements around community ownership and ethical commercialization, consistent with the Esethu framework applied to the ViXSD ASR dataset.
- BAYIBURU scripture texts: The Izzi, Ezaa, and Ikwo New Testaments (completed 1980) belong to Wycliffe Bible Translators. We do not distribute these texts or readings of them without written permission from Wycliffe. Any dataset derived from scripture readings requires that clearance first.
Datasets
Datasets in progress
| Name | Language | Kind | Size | Status | Release target | License |
|---|---|---|---|---|---|---|
| Ikwo conversational parallel corpus | Ikwo | Speech + text (parallel) | 1,000 sentences | In review | 2025 Q3 | CC BY-SA 4.0 |
| Izzi elder-speech pilot | Izzi | Speech (spontaneous) | TBD | Planned | 2026 | CC BY-SA 4.0 |
| Ezaa pilot | Ezaa | Speech (spontaneous) | TBD | Planned | 2026 | CC BY-SA 4.0 |
| Korring scoping study | Korring | Scoping | TBD | Planned | TBD | TBD |
Where to find our data
HuggingFace
Our primary distribution channel for ML-ready datasets. The org is being set up and releases will appear there first.
olubridge on HuggingFace (planned)
Zenodo
Citable, DOI-backed archival copies for academic use. Each dataset released here gets a stable identifier.
Zenodo community (planned)
ELAR
The Endangered Languages Archive at SOAS takes full deposits, including restricted items that should not appear in open ML datasets. Consent-restricted material goes here so it is preserved without being exposed.
Lanfrica
Lanfrica indexes African-language NLP resources by language and task. Our datasets will be listed there to reach researchers already looking for Ebonyi-region languages.
How to cite
When datasets carry a DOI, use that in your BibTeX entry. Until then, cite the dataset name, version, and this URL.
@misc{olubridge2025ikwo,
title = {Ikwo Conversational Parallel Corpus},
author = {{Olubridge}},
year = {2025},
note = {Version 1.0. License: CC BY-SA 4.0 with voice-cloning prohibition},
howpublished = {\url{https://olubridge.org/data}},
}A dataset-specific BibTeX entry with DOI will replace this placeholder when the first dataset is released.
Partnering on data
If you are a researcher, a language community group, a university department, or a company that wants to co-collect, co-fund, or build on Ebonyi-language data, we want to hear from you. Tell us your language, your use case, and what you can bring to the table.
We are especially interested in partnerships that put contributors at or above local prevailing rates, consistent with the pay standard set by the Thiomi dataset. If your project pays below that bar, we will talk about closing the gap before we agree.
Send a partnership enquiry