Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) non–profit founded in 2007. We make wholesale extraction, transformation, and analysis of open web data accessible to researchers.
Notes from the CommonsDB final conference in Alicante: a registry of 6.5 million open works keyed on ISCC, and the EUIPO's push for federated copyright infrastructure.