Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) non–profit founded in 2007. We make wholesale extraction, transformation, and analysis of open web data accessible to researchers.
We are publishing an experimental dataset of graph embeddings built from the Common Crawl Web Graph: a 128-dimensional vector for each of 52.9 million web hosts, learned from hyperlinks alone, along with two interactive Hugging Face Spaces.