Capabilities

Voice & Local-Language Datasets

The Hub's capability to integrate Nigerian-language speech and text data – building AI that works in Hausa, Yoruba, Igbo, and more, for inclusive national reach.

AI for Nigeria must work in Nigerian languages and voices. The Hub's capability here is integrating local-language speech and text data so that models serve the population as it actually speaks, reads, and listens – not only in English.

Why this matters for national use cases

Several identified use cases depend directly on local-language and voice data:

  • Agricultural Advisory – a multilingual chatbot over WhatsApp, Telegram, and IVR in Hausa / Yoruba / Igbo, targeting ~38 million smallholder farmers, built on the open-source N-ATLaS-LLM 8B Nigerian-language model.
  • Oral Reading Fluency (ORF) – reading-fluency assessment in local languages for ~86 million early-grade learners, using ASR (Whisper-large-v3) plus a scoring model.
  • Digital Health Coaching (Self-CAIRE) – health guidance delivered via WhatsApp and voice.

What the capability covers

  • Speech data (ASR / voice) – collecting and curating Nigerian-language audio for speech recognition and voice interfaces.
  • Text corpora – assembling local-language text for fine-tuning and evaluation.
  • Model integration – adapting open-source multilingual and Nigerian-language models to use-case domains.
  • Inclusion – ensuring dialectal and low-resource-language coverage so services reach underserved populations.

Dataset foundation

AfricanVoices platform is a concrete data layer behind this capability: 1,900 hours of curated speech and 1.9M+ sentences across Hausa, Igbo, Nigerian Pidgin, and Yoruba, with demographic metadata and customizable downloads – exactly the corpora the most use cases require.

On this page