Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) non–profit founded in 2007. We make wholesale extraction, transformation, and analysis of open web data accessible to researchers.
We made parts of our crawl archive available experimentally on a Hugging Face Storage Bucket. This blog post explains how to use the data on Hugging Face, with tools and examples on which you can build.