←  Back to Blog
September 15, 2026

CommonsDB, ISCC, Spaghetti, and Meatballs

Notes from the CommonsDB final conference in Alicante: a registry of 6.5 million open works keyed on ISCC, and the EUIPO's push for federated copyright infrastructure.
A photo of the outside of the European Union Intellectual Property Office (EUIPO), Alicante.
Outside the European Union Intellectual Property Office (EUIPO), Alicante.

I spent Monday at the final conference for CommonsDB, hosted by the EUIPO. This is a short write-up of what the project built, and what was covered at the event.

The project

CommonsDB is an EU co-funded pilot. Over 18 months it built a public registry of public domain and openly licensed works. The target was five million works. They registered almost 6.5 million.

Works are identified by ISCC (ISO 24138:2024), a code derived from the content itself, so anyone with the file can compute it. Rights status is recorded as a declaration which is in JSON-LD, signed with a key tied to a domain through a Decentralised Identifier, and timestamped. The rights statement inside is a Creative Commons or Public Domain Mark URL.

The consortium is Open Future (lead), Liccium, the Europeana Foundation, Wikimedia Sverige and IViR. It’s alongside a DG CONNECT pilot for an EU repository of public domain and open licensed works, whose purpose is to test whether such a repository improves legal certainty for reuse and gives AI developers something they can train on without a lawyer in the room.

Left-to-right: Philippe Rixhon (Valunode), Luna Schumacher (Pictoright), Krishna Sood (Black Forest Labs), and Krzysztof Nichczynski (European Commission – DG CNECT).
Left-to-right: Philippe Rixhon (Valunode), Luna Schumacher (Pictoright), Krishna Sood (Black Forest Labs), and Krzysztof Nichczynski (European Commission – DG CNECT).

Some ideas

Content-derived identifiers are glue. Europe has a lot of rights databases and no shared key between them. If the identifier can be computed from the work rather than handed out by a registrar, you can join those databases without first settling the politics of who handles all the numbers.

The idea is for this to be federated, rather than central. The EUIPO published a study in May mapping EU databases and metadata standards for copyright-protected works. Its recommendation is a federated transparency layer, with the working name “CopyrightView”, where existing providers keep authority over their own records.

The Commission put out a feasibility study in July on an EU-level registry for text and data mining opt-outs. The conclusion is that a registry could complement the existing sector-specific mechanisms, so, in other words, a meatball on top of the spaghetti of opt-out mechanisms we already have. If it ends up federated and keyed on ISCC it might be quite a useful meatball at that.

Left-to-right: Erik Schultes (Partners in FAIR), Paul Keller (Open Future), Sebastian Posth (Liccium), Thom Vaughan (Common Crawl), and Luna Schumacher (Pictoright).
Left-to-right: Erik Schultes (Partners in FAIR), Paul Keller (Open Future), Sebastian Posth (Liccium), Thom Vaughan (Common Crawl), and Luna Schumacher (Pictoright).

We are looking forward to seeing what comes of this promising pilot project.

Useful links

This release was authored by:
Thom is Principal Engineer at the Common Crawl Foundation.
Thom Vaughan
Thom is Principal Engineer at the Common Crawl Foundation.

Erratum: 

Content is truncated

Originally reported by: 
More details
Some archived content is truncated due to fetch size limits imposed during crawling. This is necessary to handle infinite or exceptionally large data streams (e.g., radio streams). Prior to March 2025 (CC-MAIN-2025-13), the truncation threshold was 1 MiB. From the March 2025 crawl onwards, this limit has been increased to 5 MiB.