Public methodology

How we record, transcribe, and archive

We publish this page because communities that share their voices deserve to know exactly what happens to them. Donors and partner institutions can audit what we do. Researchers building on our data can cite the methods directly.

Everything below describes our current practice. If we change a material detail, this page changes with it.

Collection methods

Four field-proven approaches shape how we work.

BOLD respeaking

Basic Oral Language Documentation was built for communities where writing is a barrier, and it is our primary recording workflow. A first speaker talks naturally: a story, a memory, a proverb, a description of the farm. A second speaker then listens through earphones and respeaks the same content slowly and clearly, phrase by phrase. A bilingual speaker records an oral translation into Igbo or English. No writing is needed at any step. The three-layer output (original speech, respoken version, spoken translation) makes transcription possible years later, by anyone who speaks the language.

The method is delivered through the Aikuma app, a free Android app purpose-built for this workflow. A Papua New Guinea pilot with Usarufa speakers found that participants "enjoyed recording their stories, personal narratives, songs, and dialogues after brief instruction," which documents that non-literate communities take to the app quickly. The underlying research is published in the Aikuma paper (Bird et al.).

Lig-Aikuma for parallel speech

For sessions where we want parallel speech from multiple speakers on the same prompts, we use Lig-Aikuma, an improved version of the Aikuma app developed for under-resourced language studies. It adds prompt-driven elicitation on top of the record-respeak-translate cycle and has been used with real African corpora. The method is documented in the Lig-Aikuma paper (Blachon et al.).

Community coordinator model

For volume collection, we follow the structure that produced the NaijaVoices corpus: 5,455 speakers, 1,838 hours, across roughly 18 months. The model trains local coordinators (teachers, community workers, catechists) in a single workshop, assigns each coordinator a village-level area, and pays per validated hour. Gender targets are set per session: NaijaVoices reached 58% female participation by making targets explicit at the coordinator level. We copy that structure.

Domain-anchored sessions

The Kallaama project (Gauthier et al.) in Senegal collected 125 hours across three languages by recording radio farm programmes, focus groups, and agricultural interviews rather than abstract linguistic elicitation. Elders talk at length about topics they know. We anchor our sessions in yam farming, Abakaliki rice cultivation, markets, family history, and local governance. The data quality is better, and participation rates are higher, when speakers recognize the topic as worth discussing.

Voice mark: sound waves from an open mouth

Consent stack

Every session requires two layers of consent: community-level approval from the traditional ruler or town union, followed by individual consent from each speaker. Neither replaces the other.

For non-literate participants, oral consent is standard and valid. We read the consent text aloud in the speaker's language and record the spoken agreement. ELAR accepts recorded oral consent for restricted deposits, and we follow that practice.

Any speaker can withdraw at any time. Withdrawal removes their recordings from the open portion of the corpus. Copies already exported under a restricted ELAR deposit will be flagged as withdrawn in the metadata.

The Thiomi dataset established the written precedent for this model in African speech collection. We treat their published terms as our floor, not our ceiling.

Pay and license terms

We pay at or above local daily rates. The Thiomi dataset put this principle in writing for African language data; the TIME investigation into OpenAI's Kenyan contractors (workers paid $1.32-$2/hour while the platform billed $12.50/hour per worker) documents what happens when that principle is ignored. Unclear pay is the main driver of community distrust.

Published recordings are released under CC BY-SA 4.0. Voice cloning and TTS synthesis from these recordings are prohibited. That restriction follows the precedent set by the Swivuriso ZA-ANV dataset, which was the first African speech corpus to write that prohibition explicitly into its license terms.

Community ownership is governed by the Esethu framework, which keeps downstream commercial use conditional on community agreement. Individual speakers retain attribution rights.

Metadata per recording

Dialect metadata is what made the ANV-Kenya corpus exceptional. We capture the same fields on every recording, without exception.

  • Language and variety (e.g., "Ezaa - Onueke")
  • Village and LGA
  • Speaker age band and gender
  • Consent type: oral or written, open or restricted
  • Topic domain
  • Recording date
  • Device make and model

These fields travel with every file from the moment of recording. Recordings without complete metadata are not admitted to the corpus.

Sources beyond elders

Elder speech is the foundation, but a corpus that covers only elders misrepresents how a language lives. We record across every register we can reach.

  • Market days: Eke, Orie, Afo, and Nkwo market speech, bargaining, trade numbers
  • Town criers and village announcement formulas
  • Age-grade and town-union meetings: debate and procedural speech
  • New Yam Festival: praise chants, songs, oratory
  • Folktales and moonlight stories: narrative and child-directed register
  • Proverbs sessions: elders explain each proverb (the explanation is the data)
  • Praise singers and oral poets, credited by name
  • Elders' dispute settlements and traditional courts (parties anonymized unless consent given)
  • Radio archives from Abakaliki stations, with written license
  • Local YouTube producers making Ebonyi-language video content
  • Church sermons, hymns, and live interpretations from Igbo or English into local varieties
  • Hymnals in local varieties held by parishes
Archive mark: stacked talking drums representing oral tradition

Archive plan

We archive everything twice. The open portion of each corpus release goes to HuggingFace and Zenodo for immediate research access and long-term DOI-based citation.

The full collection, including restricted recordings (sacred speech, materials held back at community request, recordings pending license clearance), goes to ELAR as a permanent deposit. ELAR is free, funder-respected, and supports tiered access controls so that restricted material stays restricted while the open portion remains freely downloadable. Funders increasingly require an ELAR deposit as a condition of grant compliance. The deposit also survives the project: if Olubridge stops operating, the recordings remain accessible to the communities and researchers who need them.