Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.

We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of June, July, and August 2025. The host-level graph consists of 691.1 million nodes and 5.0 billion edges, and the domain-level graph consists of 207.6 million nodes and 3.9 billion edges.

Laurie Burchell

Laurie is a Senior Research Engineer with Common Crawl.

Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.

Common Crawl is a 501(c)(3) non–profit founded in 2007.‍We make wholesale extraction, transformation and analysis of open web data accessible to researchers.

Over 300 billion pages spanning 15 years.

Free and open corpus since 2007.

Cited in over 10,000 research papers.

3–5 billion new pages added each month.

Featured Papers:

Lukas Kriesch, Sebastian Losacker

A geolocated dataset of German news articles

Mostafa Ansar, Anna Sperotto, Ralph Holz

Web Crawl Refusals: Insights From Common Crawl

Jeffrey Knockel, Jakub Dalek, Noura Aljizawi, Mohamed Ahmed, Levi Meletti, and Justin Lau

Banned Books: Analysis of Censorship on Amazon.com

Xian Gong, Paul X. McCarthy, Marian-Andrei Rizoiu, Paolo Boldi

Harmony in the Australian Domain Space

Kevin Saric, Felix Savins, Gowri Sankar Ramachandran, Raja Jurdak, Surya Nepal

Hyperlink Hijacking: Exploiting Erroneous URL Links to Phantom Domains

Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Mingchuan Zhang, Y.K. Li, Y. Wu, Daya Guo

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Asier Gutiérrez-Fandiño, David Pérez-Fernández, Jordi Armengol-Estapé, David Griol, Zoraida Callejas

esCorpius: A Massive Spanish Crawling Corpus

Marius Løvold Jørgensen, UiT Norges Arktiske Universitet

BacklinkDB: A Purpose-Built Backlink Database Management System

Latest Blog Post:

Host- and Domain-Level Web Graphs June, July, and August 2025

The Data

Resources

Community

About

Common Crawl is a 501(c)(3) non–profit founded in 2007.
‍
We make wholesale extraction, transformation and analysis of open web data accessible to researchers.