Blog

The latest news, interviews, technologies, and resources.

Filter by Category or Search by Title

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Video Tutorial: MapReduce for the Masses

Video Tutorial: MapReduce for the Masses

Learn how you can harness the power of MapReduce data analysis against the Common Crawl dataset with nothing more than five minutes of your time, a bit of local configuration, and 25 cents.
Gil Elbaz and Nova Spivack on This Week in Startups

Gil Elbaz and Nova Spivack on This Week in Startups

Nova and Gil, in discussion with host Jason Calacanis, explore in depth what Common Crawl is all about and how it fits into the larger picture of online search and indexing. Underlying their conversation is an exploration of how Common Crawl's open crawl of the web is a powerful asset for educators, researchers, and entrepreneurs.
Video: This Week in Startups - Gil Elbaz and Nova Spivack

Video: This Week in Startups - Gil Elbaz and Nova Spivack

Nova and Gil, in discussion with host Jason Calacanis, explore in depth what Common Crawl is all about and how it fits into the larger picture of online search and indexing.
MapReduce for the Masses: Zero to Hadoop in Five Minutes with Common Crawl

MapReduce for the Masses: Zero to Hadoop in Five Minutes with Common Crawl

Common Crawl aims to change the big data game with our repository of over 40 terabytes of high-quality web crawl information into the Amazon cloud, the net total of 5 billion crawled pages.
Common Crawl Discussion List

Common Crawl Discussion List

We have started a Common Crawl discussion list to enable discussions and encourage collaboration between the community of coders, hackers, data scientists, developers and organizations interested in working with open web crawl data.
Answers to Recent Community Questions

Answers to Recent Community Questions

In this post we respond to the most common questions. Thanks for all the support and please keep the questions coming!
Common Crawl Enters A New Phase

Common Crawl Enters A New Phase

A little under four years ago, Gil Elbaz formed the Common Crawl Foundation. He was driven by a desire to ensure a truly open web. He knew that decreasing storage and bandwidth costs, along with the increasing ease of crunching big data, made building and maintaining an open repository of web crawl data feasible.
Video: Gil Elbaz at Web 2.0 Summit 2011

Video: Gil Elbaz at Web 2.0 Summit 2011

Hear Common Crawl founder discuss how data accessibility is crucial to increasing rates of innovation as well as give ideas on how to facilitate increased access to data.

Common Crawl Blog

Measuring Crawled Coverage of a Website in Common Crawl

July 20, 2026

How can we measure how many pages we’ve crawled from a particular website? The answer is a lot more complicated than you might think.

Read More...

Is one vantage point enough? IPv6 across the top million web hosts

July 17, 2026

We probed the top 1,000,000 web hosts for IPv6 from five vantage points on three continents. 31.7% work from everywhere, one vantage point turns out to be enough for the headline rate, and 5,530 hosts reveal why it isn't enough for the rest of the story.

Read More...

Turning 30,000 Arabic Domains Into a Better Crawl

July 3, 2026

How we filtered, geolocated and categorised a donation of Arabic seed domains

Read More...

Common Crawl Foundation at LREC 2026

June 30, 2026

The Common Crawl team attended the 16th International Conference on Language Resources and Evaluation in Palma, Mallorca, co-organizing a tutorial, presenting recent published work, and strengthening links with the research community.

Read More...

13th Web-as-Corpus Workshop @ EMNLP 2026

June 29, 2026

The WaC-13 workshop invites research submissions on web data, corpus building, and linguistic analysis.

Read More...

Host- and Domain-Level Web Graphs April, May, and June 2026

June 25, 2026

We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of April, May, and June 2026. The graphs consist of 247.3 million nodes and 6.3 billion edges at the host level, and 121.1 million nodes and 3.9 billion edges at the domain level.

Read More...

June 2026 Crawl Archive Now Available

June 22, 2026

We are happy to announce the release of the June 2026 crawl archive, consisting of 2.10 billion web pages, or 354.59 TiB of uncompressed content.

Read More...

CommonLID Update: New Tools, Growing Impact

June 16, 2026

CommonLID, a community-built language ID benchmark, has a new website and interactive leaderboard. Its paper was accepted to ACL 2026, with a poster session on 7 July. Source code, a PyPI package, and the dataset are now available.

Read More...

Common Crawl Foundation at IIPC-WAC 2026

June 10, 2026

Common Crawl was well represented with contributions at the 2026 IIPC Web Archiving Conference and General Assembly.

Read More...

The Columnar Index Is Now the URL Index

June 3, 2026

We have renamed the Columnar Index to the URL Index, to be clearer about its purpose and to pave the way for more datasets in a columnar format.

Read More...

Introducing the AI Visibility Audit

June 1, 2026

A free guide for SEOs and GEOs on how to check whether AI systems can actually reach a site, and how to stay visible in the crawl that trains them.

Read More...

Host- and Domain-Level Web Graphs March, April, and May 2026

May 29, 2026

We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of March, April, and May 2026. The graphs consist of 262.4 million nodes and 8.1 billion edges at the host level, and 118.8 million nodes and 4.3 billion edges at the domain level.

Read More...

May 2026 Crawl Archive Now Available

May 25, 2026

We are happy to announce the release of the May 2026 crawl archive, consisting of 2.16 billion web pages, or 365.56 TiB of uncompressed content.

Read More...

April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket

May 20, 2026

As an early experiment in distributing Common Crawl data through another channel, the April 2026 crawl archive is now available in a Hugging Face Storage Bucket, alongside its existing home on AWS S3.

Read More...

You can now build directly on Common Crawl from the browser

May 6, 2026

Browsers can now fetch Common Crawl data directly, no backend needed. Build SQL explorers, snapshot viewers and diff tools as static pages.

Read More...

Host- and Domain-Level Web Graphs February, March, and April 2026

April 30, 2026

We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of February, March, and April 2026. The graphs consist of 269.0 million nodes and 9.4 billion edges at the host level, and 124.6 million nodes and 4.8 billion edges at the domain level.

Read More...

April 2026 Crawl Archive Now Available

April 28, 2026

We are pleased to announce that the crawl archive for April 2026 is now available, containing 2.19 billion web pages or 379.2 TiB of uncompressed content.

Read More...

April 2026 Common Crawl Newsletter

April 6, 2026

Check out our newsletter for April 2026, with updates on what we've been up to.

Read More...

Announcing a Change to Common Crawl Dataset Size Reporting

April 1, 2026

Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit to fall out of use. Our latest crawl now exceeds 689 tebibbles.

Read More...

Host- and Domain-Level Web Graphs January, February, and March 2026

March 24, 2026

We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of January, February, and March 2026. The graphs consist of 270.2 million nodes and 9 billion edges at the host level, and 120 million nodes and 4.4 billion edges at the domain level.

Read More...

March 2026 Crawl Archive Now Available

March 19, 2026

We are pleased to announce the release of the March 2026 crawl, containing 1.97 billion web pages, or 344.64 TiB of uncompressed content. We also observed a dramatic increase in fetches over IPv6, explained by the enabling of Happy Eyeballs in the OkHttp library.

Read More...

IPv6 Adoption Across the Top 100K Web Hosts

March 16, 2026

We probed the 100,000 most-linked web hosts for IPv6 support using the Common Crawl Web Graph. Only 36.9% are fully reachable over IPv6, with adoption ranging from 71% among the top 100 to 32% in the long tail.

Read More...

Web Graph Statistics Gets a Proper Upgrade

March 6, 2026

Our Web Graph Statistics site has been updated with interactive charts, a domain lookup tool for tracking harmonic centrality and PageRank over time, mobile improvements, unified rank tables with OR filtering, and merged degree plots.

Read More...

Measuring Web Accessibility from Crawl Archives

March 2, 2026

A WCAG colour contrast audit of 240 top domains using Common Crawl's February 2026 archive finds four in ten colour pairings fall short of accessibility thresholds. Only one in five sites are fully compliant.

Read More...

Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java

February 26, 2026

Introducing the second installment in our Whirlwind Tour series, covering crawl structure, index access, and content extraction, giving developers a practical foundation for building Java-based data workflows.

Read More...