ProjectsProjects Archive

Masakhane

Masakhane – the grassroots pan-African NLP collective building open datasets, models, and benchmarks for African languages, including Nigerian languages and health-relevant resources directly reusable by NAISH.

Masakhane ("we build together" in isiZulu) is a grassroots, pan-African research collective strengthening natural-language processing for African languages – "for Africans, by Africans." Since ~2019 it has grown from a machine-translation effort into a community producing open datasets, models, and benchmarks across translation, named-entity recognition, news classification, question answering, part-of-speech tagging, and speech. It is included in this archive because its artifacts already cover the major Nigerian languages and include health-relevant resources that NAISH language use cases can reuse directly rather than rebuild.

Reference: masakhane.io

Participatory, community-owned. Masakhane works openly (Slack, weekly meetings, Colab notebooks) through "participatory research" and "data archeology" – co-creating datasets and models with researchers, linguists, and language communities. Its model for ethically sourcing local-language data is a reusable template for Nigerian health-language work.

Scale

1,000+
Contributors
30+
Countries
38+
African languages

Figures are self-reported by masakhane.io and grow over time. Masakhane's collaborative papers appear at top venues (ACL, EMNLP, NAACL, ICLR AfricaNLP), and its Masakhane MT web interface received a 2021 Wikimedia Foundation Research Award.

Projects

Datasets & artifacts

Masakhane publishes openly on Hugging Face and GitHub – hundreds of models and datasets, most built on African-language pretrained models such as AfroXLMR.

Key papers: Masakhane – Machine Translation for Africa · Participatory Research for Low-resourced MT · MasakhaNER · MasakhaNER 2.0 · MasakhaNEWS · MAFAND-MT · AfriQA · MasakhaPOS · AfroBench: How good are LLMs on African languages?

Relevance to NAISH

  • Nigerian-language coverage – Hausa, Yoruba, Igbo, and Nigerian Pidgin appear across MAFAND-MT, MasakhaNER, MasakhaNEWS, and AfriQA – the exact languages needed for frontline health communication in Nigeria.
  • Health-specific artifacts – MasakhaNEWS includes a health topic category and AfrIFact spans healthcare content, useful for building trustworthy health QA and combating health misinformation in local languages.
  • Low-resource MT & QA foundations – MAFAND-MT and AfriQA provide parallel data and cross-lingual QA to translate health guidance and answer patient / community-health-worker questions.
  • Speech / ASR – the Masakhane African Languages Hub (co-funded by the Gates Foundation, Google.org, and others) is building large-scale, culturally-grounded ASR/voice datasets for African languages, explicitly citing speech-based health services – a strong strategic and funding alignment for NAISH.

Rather than commission new Nigerian-language corpora from scratch, NAISH language use cases can adopt Masakhane's open datasets, models, and participatory data model – and partner through the shared Gates-funded Hub.

On this page