Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) non–profit founded in 2007. We make wholesale extraction, transformation and analysis of open web data accessible to researchers.
Two weeks in Basel and Vienna, at the HTTP Workshop and IETF 126. Protocol adoption measured across the whole web, and an attempt to define what "machine readable" actually means.