< Back to Blog
January 8, 2014

Winter 2013 Crawl Data Now Available

The second crawl of 2013 is now available! In late November, we published the data from the first crawl of 2013. The new dataset was collected at the end of 2013, contains approximately 2.3 billion webpages and is 148TB in size.
Common Crawl Foundation
Common Crawl Foundation
Common Crawl - Open Source Web Crawling data‍

The second crawl of 2013 is now available! In late November, we published the data from the first crawl of 2013 (see previous blog post for more detail on that dataset). The new dataset was collected at the end of 2013, contains approximately 2.3 billion webpages and is 148TB in size. The new data is located in the commoncrawl bucket at /crawl-data/CC-MAIN-2013-48/

Data Type File List #Files Total Size
Compressed (TiB)
Segments segment.paths.gz 519
WARC warc.paths.gz 51900 31.93
WAT wat.paths.gz 45195 9.64
WET wet.paths.gz 45195 3.36
URL index files cc-index.paths.gz 302 0.14
Columnar URL index files cc-index-table.paths.gz 300 0.15

In 2013, we made changes to our crawling and post-processing systems. As detailed in the previous blog post, we switched file formats to the international standard WARC and WAT files. We also began using Apache Nutch to crawl – stay tuned for an upcoming blog post on our use of Nutch. The new crawling method relies heavily on the generous data donations from blekko and we are extremely grateful forongoing support!

In 2014 we plan to crawl much more frequently and publish fresh datasets at least once a month.

Errata
No items found.
This release was authored by:
No items found.