Building Speech Technology in Indigenous Languages with Mozilla Common Voice
🎥 Session recording: [https://youtu.be/u6tMzowZ10M?t=1264](https://youtu.be/u6tMzowZ10M?t=1264)
This hands-on session explores the development of Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models for Indigenous languages. It will provide an in-depth overview of how we leverage Mozilla Common Voice for large-scale speech data collection, validation, and annotation to support the creation of open, community-driven language technologies for underrepresented languages.
Participants will gain practical insights into the distinctive linguistic characteristics of many Indigenous languages such as tonal systems, vowel harmony, and rich morphological structures and how these features influence data preparation and model training. The session will also examine both the technical and community-related challenges involved in developing robust ASR and TTS systems for low-resource contexts, as well as the significant opportunities this work presents for advancing digital inclusion.
By the end of the session, participants will better understand how structured community contributions, quality assurance processes, and open collaboration can accelerate progress in building reliable, culturally grounded speech technologies that reflect the voices and knowledge systems of Indigenous language communities.
<br>