Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) non–profit founded in 2007. We make wholesale extraction, transformation and analysis of open web data accessible to researchers.
We are pleased to announce a new release of host-level and domain-level Web Graphs based on the crawls of February, March, and April 2025. The graph consists of 309.2 million nodes and 2.9 billion edges at the host level, and 157.1 million nodes and 2.1 billion edges at the domain level.
Thom Vaughan
Thom is Principal Technologist at the Common Crawl Foundation.