Building Open Datasets for Global Majority Languages
The proposed discussion aims to gather a set of best practices related to language documentation and creation of datasets for the development of AI models. Data collection for building AI technologies has largely been undertaken by a specific set of stakeholders invested in developing LLMs; open knowledge projects like Wikimedia platforms are a major source of such data. However, the data available in various languages, particularly underrepresented and largely oral languages, is still scarce, and faces issues related to quality and curation. Another challenge is the lack of digitised, domain-specific datasets that contribute to the development of smaller, sovereign AI models. These would strengthen open knowledge projects through community-led dataset co-creation and fair, ethical ownership and use. This discussion may therefore be replicated across various languages and Wikimedia projects, to gather insights on language documentation and the forms of data. Importantly, this model may also promote data minimization and contribute to building a more robust data commons in the long-term. Efforts in the Wikimedia and the larger open knowledge movement already offer good examples, such as Lingua Libre, and Mozilla’s Common Voice. For example, Common Voice offers a model for co-created voice datasets in over 290 languages, which may then be used for a wide range of functionalities. While these projects may not have been developed with their potential use for training of AI models, they now offer a way to imagine co-creation of open datasets, with diversity, equity and access by the commons.