Masakhane
Masakhane – the grassroots pan-African NLP collective building open datasets, models, and benchmarks for African languages, including Nigerian languages and health-relevant resources directly reusable by NAISH.
Masakhane ("we build together" in isiZulu) is a grassroots, pan-African research collective strengthening natural-language processing for African languages – "for Africans, by Africans." Since ~2019 it has grown from a machine-translation effort into a community producing open datasets, models, and benchmarks across translation, named-entity recognition, news classification, question answering, part-of-speech tagging, and speech. It is included in this archive because its artifacts already cover the major Nigerian languages and include health-relevant resources that NAISH language use cases can reuse directly rather than rebuild.
Reference: masakhane.io
Participatory, community-owned. Masakhane works openly (Slack, weekly meetings, Colab notebooks) through "participatory research" and "data archeology" – co-creating datasets and models with researchers, linguists, and language communities. Its model for ethically sourcing local-language data is a reusable template for Nigerian health-language work.
Scale
Figures are self-reported by masakhane.io and grow over time. Masakhane's collaborative papers appear at top venues (ACL, EMNLP, NAACL, ICLR AfricaNLP), and its Masakhane MT web interface received a 2021 Wikimedia Foundation Research Award.
Projects
Masakhane MT / MMT
The founding effort (ICLR AfricaNLP 2020): community-built neural machine-translation models and benchmarks across dozens of African languages.
Decolonise Science
Machine translation of scientific text into African languages so science can be discussed in indigenous languages.
MasakhaNER
'Know Our Names' – the largest human-annotated African NER corpus; MasakhaNER 2.0 spans 20 languages.
MasakhaNEWS
News-topic classification across 16 African languages, including a health category.
MAFAND-MT
Human-translated news-domain parallel MT dataset for 16 African languages.
AfriQA
First cross-lingual open-retrieval QA dataset for African languages – 12,000+ examples across 10 languages.
MasakhaPOS
Part-of-speech tagging for typologically diverse African languages.
Datasets & artifacts
Masakhane publishes openly on Hugging Face and GitHub – hundreds of models and datasets, most built on African-language pretrained models such as AfroXLMR.
- Hugging Face org – huggingface.co/masakhane
- masakhane/masakhanews – news classification (incl. health category)
- masakhane/afriqa – cross-lingual QA
- AfriMMLU / AfriMGSM – multilingual understanding & math benchmarks
- AfrIFact – multilingual retrieval & fact-checking benchmark explicitly covering healthcare content
- AfriCOMET / AfriMTE – MT evaluation metrics for African languages
- GitHub org – github.com/masakhane-io (masakhane-mt, masakhane-ner, lafand-mt, afriqa, masakhane-pos)
Key papers: Masakhane – Machine Translation for Africa · Participatory Research for Low-resourced MT · MasakhaNER · MasakhaNER 2.0 · MasakhaNEWS · MAFAND-MT · AfriQA · MasakhaPOS · AfroBench: How good are LLMs on African languages?
Relevance to NAISH
- Nigerian-language coverage – Hausa, Yoruba, Igbo, and Nigerian Pidgin appear across MAFAND-MT, MasakhaNER, MasakhaNEWS, and AfriQA – the exact languages needed for frontline health communication in Nigeria.
- Health-specific artifacts – MasakhaNEWS includes a health topic category and AfrIFact spans healthcare content, useful for building trustworthy health QA and combating health misinformation in local languages.
- Low-resource MT & QA foundations – MAFAND-MT and AfriQA provide parallel data and cross-lingual QA to translate health guidance and answer patient / community-health-worker questions.
- Speech / ASR – the Masakhane African Languages Hub (co-funded by the Gates Foundation, Google.org, and others) is building large-scale, culturally-grounded ASR/voice datasets for African languages, explicitly citing speech-based health services – a strong strategic and funding alignment for NAISH.
Rather than commission new Nigerian-language corpora from scratch, NAISH language use cases can adopt Masakhane's open datasets, models, and participatory data model – and partner through the shared Gates-funded Hub.
Links
Dimagi
Dimagi – the social enterprise behind CommCare and Open Chat Studio, a long-standing Gates Foundation partner whose open-source, offline-first, LLM-enabled platforms for frontline health and development work map directly onto several NAISH use cases.
Gooey.AI
Gooey.AI – a low-code platform for building and deploying multilingual AI copilots over WhatsApp, voice, SMS, and web, with proven frontline health and agriculture deployments (Farmer Chat, UN IOM Afiya) directly relevant to NAISH delivery.
