Governing the Open Web in the Age of AI: From Legal Principles to Technical Practice
As AI systems increasingly rely on web data, the governance of the open web is entering a period of rapid change. Legal frameworks, technical standards, and community norms are evolving simultaneously, raising fundamental questions about who can access knowledge, under what conditions, and with what safeguards.
This workshop brings together legal and technical expertise from the Wikimedia movement and the Common Crawl Foundation to connect principles with practice.
We begin with an overview of the legal frameworks and recent AI and copyright cases, such as LAION or OpenAI vs. New York Times, including text and data mining exceptions, licensing, and ongoing policy debates on opt-out. It will give an overview of how public interests are currently interpreted in different jurisdictions by taking into account the meaning for the Wikimedia movement and CC-licences.
In the second part of this workshop, we examine how these principles are implemented in practice at Common Crawl. Specifically, we will introduce responsible web crawling practices, including adherence to robot.txt, crawler politeness, transparent identification, and opt-out processes. The Wikimedia movement will gain insight into how public web archives are created and maintained as shared infrastructure for research and open knowledge.
At the end, we want to discuss together with participants how the growing use of crawler blocking, the proliferation of AI-specific opt-outs, and their implications for the accessibility and diversity of web data, can change the open data governance in the age of AI.