From Abstract Content to Concrete Text with Wikidata Lexemes
Projects like Abstract Wikipedia will only succeed with thorough, extensible systems for generating natural encyclopedic language from a single source of informational content. While the current state of ‘language generation’ on Abstract Wikipedia lacks sufficient thoroughness and extensibility, such limitations do not prevent building other systems that use well-developed Wikidata lexemes and items to produce text.
This workshop explores how the Ninai/Udiron system (developed by the presenter) takes Wikidata entities from across diverse languages and makes natural text generation possible in those languages, with help from linguistic references and user preferences. Participants will explore how to write abstract content that expresses complex relationships, represented by both Wikidata statements and assertions too complex for such statements, and how to indicate different stylistic choices in that content. Participants will also learn how to indicate that particular words and expressions should appear under these stylistic choices, and to set up terms that only appear in their language so that they may be imported into other languages. The French and Breton languages will serve as primary pivots around which discussion of other languages’ lexemes and behaviors, including around how concepts may be re-expressed through processes such as loan translation, etymological equivalency, and transliteration.
The presenter hopes that the workshop empowers participants to improve at least one aspect of Wikidata’s lexicographical data in their own languages, so that readers can ultimately benefit from greater possibilities regarding the expression of both encyclopedic and non-encyclopedic information across all languages.