Expanding the Language and Cultural Coverage of Common Crawl
December 11, 2024
We aim to enhance linguistic diversity in our dataset by inviting community contributions of non-English URLs and collaborating with MLCommons on a Language Identification campaign.
Read More...October/November 2024 Newsletter
November 25, 2024
We’re pleased to announce this month's newsletter, featuring key updates, upcoming events, and community highlights.
Read More...Host- and Domain-Level Web Graphs September, October, November 2024
November 20, 2024
We are pleased to announce a new release of host-level and domain-level Web Graphs based on the crawls of September, October, and November 2024. The crawls used to generate the graphs were CC-MAIN-2024-46, CC-MAIN-2024-42, and CC-MAIN-2024-38.
Read More...November 2024 Crawl Archive Now Available
November 18, 2024
The crawl archive for November 2024 is now available. The data was crawled between November 1st and November 15th, and contains 2.68 billion web pages (or 405 TiB of uncompressed content). Page captures are from 47.5 million hosts or 38.3 million registered domains and include 1 billion new URLs, not visited in any of our prior crawls.
Read More...Reflections on Recent Talks at the Turing Institute and UCL
November 4, 2024
Thom Vaughan and Pedro Ortiz Suarez discussed the power of Common Crawl’s open web data in driving research and innovation during two notable presentations last week.
Read More...Introducing the Common Crawl Errata Page for Data Transparency
October 30, 2024
As part of our commitment to accuracy and transparency, we are pleased to introduce a new Errata page on our website.
Read More...Host- and Domain-Level Web Graphs August, September, and October 2024
October 22, 2024
We are pleased to announce a new release of host-level and domain-level Web Graphs based on the crawls of August, September, and October 2024. The crawls used to generate the graphs were CC-MAIN-2024-33, CC-MAIN-2024-38, and CC-MAIN-2024-42.
Read More...October 2024 Crawl Archive Now Available
October 20, 2024
The data was crawled between October 3rd and October 16th, and contains 2.49 billion web pages (or 365 TiB of uncompressed content). Page captures are from 47.5 million hosts or 38.3 million registered domains and include 1.03 billion new URLs, not visited in any of our prior crawls.
Read More...White House Briefing on Open Data’s Role in Technology
October 8, 2024
We recently had the honor of briefing the White House Office of Science and Technology Policy (OSTP) on the role of The Common Crawl Foundation as critical infrastructure in the artificial intelligence ecosystem and how we can support U.S. federal efforts in advancing responsible AI use and research.
Read More...IAB Workshop on AI-CONTROL
September 30, 2024
Earlier this month, the Common Crawl Foundation had the privilege of participating in a groundbreaking workshop hosted by the Internet Architecture Board (IAB) in Washington DC.
Read More...Host- and Domain-Level Web Graphs July, August, and September 2024
September 26, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of July, August, and September 2024. The crawls used to generate the graphs were CC-MAIN-2024-30, CC-MAIN-2024-33, and CC-MAIN-2024-38.
Read More...September 2024 Crawl Archive Now Available
September 24, 2024
The crawl archive for September 2024 is now available. The data was crawled between September 7th and September 21st 2024, and contains 2.8 billion web pages (or 410 TiB of uncompressed content).
Read More...August/September 2024 Newsletter
September 10, 2024
We're pleased to announce our newsletter for August and September 2024.
Read More...Host- and Domain-Level Web Graphs June, July, and August 2024
August 21, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of June, July, August 2024. The crawls used to generate the graphs were CC-MAIN-2024-33, CC-MAIN-2024-30, and CC-MAIN-2024-26.
Read More...August 2024 Crawl Archive Now Available
August 18, 2024
The crawl archive for August 2024 is now available. The data was crawled between August 3rd and August 16th, and contains 2.3 billion web pages (or 327.4 TiB of uncompressed content).
Read More...The Increase of Common Crawl Citations in Academic Research
August 6, 2024
Common Crawl's impact on research has grown substantially since its beginning. Our crawls have become a vital resource for researchers in various fields, from natural language processing to red teaming.
Read More...Host- and Domain-Level Web Graphs May, June, and July 2024
July 30, 2024
We are pleased to announce a new release of host-level and domain-level Web Graphs based on the crawls of May, June, and July 2024.
Read More...July 2024 Crawl Archive Now Available
July 28, 2024
We are pleased to announce that the crawl archive for July 2024 is now available, containing 2.5 billion web pages, or 360 TiB of uncompressed content.
Read More...Common Crawl Statistics Now Available on Hugging Face
July 22, 2024
We're excited to announce that Common Crawl’s statistics are now available on Hugging Face!
Read More...The Environmental Impact of the Cloud - the Common Crawl Case Study
July 16, 2024
Looking at tools (Green Software) and methodologies to evaluate the environmental impact of the cloud (a nascent activity coined GreenOps).
Read More...Host- and Domain-Level Web Graphs April, May, and June 2024
June 30, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of April, May, June 2024. The crawls used to generate the graphs were CC-MAIN-2024-18, CC-MAIN-2024-22, and CC-MAIN-2024-26.
Read More...Dialog and Discovery at AI_dev 2024
June 28, 2024
This month members from the Common Crawl Foundation attended the AI_dev: Open Source GenAI & ML Summit in Paris, where discussions focused on AI advancements, ethics, and Open Source solutions.
Read More...June 2024 Crawl Archive Now Available
June 28, 2024
The crawl archive for June 2024 is now available. The data was crawled between June 12th and June 26th, and contains 2.7 billion web pages (or 382 TiB of uncompressed content). Page captures are from 52.7 million hosts or 41.4 million registered domains and include 945 million new URLs, not visited in any of our prior crawls.
Read More...May/June 2024 Newsletter
June 25, 2024
We’re pleased to share our newsletter for May/June 2024, featuring the latest updates, events, and highlights from our community.
Read More...Host- and Domain-Level Web Graphs February/March, April, and May 2024
June 4, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of February, April, and May 2024.
Read More...