Papers
arxiv:2602.02734

WAXAL: A Large-Scale Multilingual African Language Speech Corpus

Published on Feb 2
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

WAXAL is a large-scale speech dataset for 21 Sub-Saharan African languages providing ASR and TTS resources to address the digital divide in low-resource languages.

The advancement of speech technology has predominantly favored high-resource languages, creating a significant digital divide for speakers of most Sub-Saharan African languages. To address this gap, we introduce WAXAL, a large-scale, openly accessible speech dataset for 21 languages representing over 100 million speakers. The collection consists of two main components: an Automated Speech Recognition (ASR) dataset containing approximately 1,250 hours of transcribed, natural speech from a diverse range of speakers, and a Text-to-Speech (TTS) dataset with over 180 hours of high-quality, single-speaker recordings reading phonetically balanced scripts. This paper details our methodology for data collection, annotation, and quality control, which involved partnerships with four African academic and community organizations. We provide a detailed statistical overview of the dataset and discuss its potential limitations and ethical considerations. The WAXAL datasets are released at https://huggingface.co/datasets/google/WaxalNLP under the permissive CC-BY-4.0 license to catalyze research, enable the development of inclusive technologies, and serve as a vital resource for the digital preservation of these languages.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2602.02734
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 23

Browse 23 models citing this paper

Datasets citing this paper 18

Browse 18 datasets citing this paper

Spaces citing this paper 10

Browse 10 spaces citing this paper

Collections including this paper 3