Wikimania 2026

Who's Missing? Using AI to Spot Wikipedia's Coverage Gaps

Session type: Lightning talk Showcase
Track: Artificial intelligence

Speakers

Jonathan Deamer

I've been editing English Wikipedia for over 15 years. In 2024 I had the privilege of attending my first Wikimania.

Over the last couple of years I've developed a particular interest in writing biographies of British artists, focusing in particular on winners of the John Moores Painting Prize, based in my adopted city of Liverpool.

My career has largely been at the intersection of technology and media, previously at a large book publisher and now as a data consultant at US-based AI company.
Outside of Wikimedia activities, I'm an enthusiastic runner, music trivia anorak, and I always travel with my Aeropress coffee maker.

Abstract

With so much news published every day, how can editors keep on top of new information that could enable additions to Wikipedia? Many notable people are covered in reliable sources but never become Wikipedia articles because no editor happens to notice them. Wikipedia's coverage gaps are often a discovery problem, not a sourcing problem.

In this talk, I'll describe an AI tool I've built to help solve this problem. The tool automates the monitoring of news sources, researches content gaps and signs of notability, and produces a shortlist of people worth looking into. Editorial judgment and article creation are then an entirely human process, with no automated editing.

This is a reusable approach: the AI handles the tedious monitoring part, while humans decide what's actually worth writing about. I have already used this process to write articles that have gone on to be featured on English Wikipedia's Main Page.

I will outline how editors can adapt this approach using their own chosen sources, and how individual experiments could inspire shared tooling that highlights potential new Wikipedia articles as soon as reliable sources become available.

This begins to address one of the core questions around the use of AI in Wikimedia projects: how can it make contributors' lives easier, and further goals like knowledge equity, without undermining the core principles that make Wikimedia projects trustworthy, human and fun?

Additional information

How does your session relate to the event theme: Liberté, Équité, Fiabilité (Freedom, Equity, Reliability).
The core aim of this talk is to demonstrate how AI can strengthen representation on Wikipedia (equity), without undermining core principles that make it trustworthy (reliability), and in a way that's accessible and reusable by anyone (freedom). Freedom: This approach doesn't dictate a particular way of editing or creating content and can be tailored to use editor-selected news sources, allowing contributors to focus on the topics, regions, or communities they care about. An example implementation has been released as open source software, so it's open to community improvement. Equity: Many missing biographies of notable people aren't because of a lack of reliable sources. This workflow allows editors to monitor diverse or locally relevant sources, and so creates fairer opportunities for notable individuals to be surfaced and considered for inclusion. Reliability: This model starts with independent, reliable sources and topics selected by a human editor. The AI finds people who keep appearing in quality sources; I check whether they're actually notable and whether the sourcing holds up. This AI-assisted but human led approach is counter to some of the common fears about what AI could mean for the encyclopedia.
Which Wikimedia audiences will find this content the most useful?
This session will be most useful to: - Active editors who create new articles and want a more systematic way to discover well-sourced topics. - Contributors involved in addressing content gaps, particularly around underrepresented biographies. - Technically curious Wikimedians interested in responsible, human-in-the-loop AI workflows. The session will avoid technical implementation details and instead focus on the repeatable approach and its guardrails. Participants who want to explore further will be directed to the open-source implementation and documentation after the session. The goal is not to teach a tool in depth during the talk, but to introduce a practical, reproducible idea that editors can adapt.
What is the experience level needed for the audience for your session?
Average knowledge about Wikimedia projects or activities