The crawl archive for September 2018 is now available! It contains 2.8 billion web pages and 220 TiB of uncompressed content, crawled between September 17th and 26th.
The following improvements and fixes to the data formats have been made:
- the columnar index contains the content language of a web page as a new field. Please read the instructions below how to upgrade your tools to read newly added fields.
- WARC revisit records (HTTP status 304) in the URL indexes do not include a field for the payload "digest" anymore. The corresponding column "content_digest" in the columnar index now contains null values.
- we've fixed a bug in the WARC writer which added an extra line break (\r\n) between HTTP header and payload in WARC response record. See the announcement on our Google group for details. Thanks again to Greg Lindahl for discovering this bug!
The September crawl contains 500 million new URLs, not contained in any crawl archive before. New URLs stem from
- the continued seed donation of URLs from mixnode.com
- extracting and sampling URLs from sitemaps, RSS and Atom feeds if provided by hosts visited in prior crawls. Hosts are selected from the highest-ranking 60 million domains of the May/June/July 2018 webgraph data set
- a breadth-first side crawl within a maximum of 6 links (“hops”) away from the home pages of the top 25 million domains of the webgraph dataset
- a random sample taken from WAT files of the August crawl
New Fields in the Columnar URL Index
The columnar index has been updated to contain two new fields added to WARC and CDX files starting with the August crawl:
- content_charset: the character encoding used by the HTML page
- content_languages: a comma-separated list of ISO-639-3 language codes detected identified by the Compact Language Detector 2 (CLD2)
In addition, the column content_digest now contains null values.
The table schema in the cc-index-table project on github has been updated to reflect these changes.
Please follow the instructions below to upgrade to the new schema for Spark, Athena/Presto or Hive. If you do not want to use the new fields, no action is required, the tools should continue to work with the old schema.
Spark
The property spark.sql.parquet.mergeSchema must be set to true, e.g. by running the Spark job with the command
Note that enabling schema merging has a negative impact on the performance of Spark jobs, you may want to enable it only in case the new fields are required for your task.
Athena / Presto
Please create a new table using the updated schema. The old schema will continue to work but the new fields cannot be used. Further information can be found in the chapter about schema updates in the Athena documentation.
Hive
Schema evolution is supported since version 0.13. The procedure is essentially the same as for Athena – you need to drop and re-create the table with the updated schema only in case the new fields are used.
Archive Location and Download
The September crawl archive is located in the commoncrawl bucket at crawl-data/CC-MAIN-2018-39/.
To assist with exploring and using the dataset, we provide gzipped files which list all segments, WARC, WAT and WET files.
By simply adding either s3://commoncrawl/ or https://data.commoncrawl.org/ to each line, you end up with the S3 and HTTP paths respectively.
The Common Crawl URL Index for this crawl is available at: https://index.commoncrawl.org/CC-MAIN-2018-39/. Also the columnar index has been updated to contain this crawl.
We are grateful to our friends at mixnode for donating a seed list of 200 Million URLs to enhance the Common Crawl.
Please donate to Common Crawl if you appreciate our free datasets! We’re also seeking corporate sponsors to partner with Common Crawl for our non-profit work in open data. Please contact info@commoncrawl.org for sponsorship information.
Erratum:
WAT data: repeated WARC and HTTP headers are not preserved
Repeated HTTP
and WARC
headers were not represented in the JSON
data in WAT
files. When a header was repeated adding a further value of that header, only the last value was stored and other values were lost. This issues was fixed with CC-MAIN-2024-51
, see ia-web-commons#18. All WAT
files from CC-MAIN-2013-20
until CC-MAIN-2024-46
are affected.
Erratum:
WARC revisit metadata records
The revisit records in the Common Crawl WARC
archives in all crawls from CC-MAIN-2018-34
to CC-MAIN-2024-46
(since Aug 2018) lack the metadata record which is attached to all response records. Fixed with CC-MAIN-2024-51
, see commoncrawl/nutch#33. Note: before CC-MAIN-2018-34
, WARC
revisit records were not stored at all.
Erratum:
Erroneous title field in WAT records
The "Title" extracted in WAT records to the JSON path `Envelope > Payload-Metadata > HTTP-Response-Metadata > HTML-Metadata > Head > Title
` is not the content included in the <title>
element in the HTML header (<head>
element) if the page contains further <title>
elements in the page body. The content of the last <title>
element is written to the WAT "Title". This bug was observed if the HTML page includes embedded SVG graphics.
The issue was reported by the user Robert Waksmunski:
- https://groups.google.com/g/common-crawl/c/ZrPFdY3pPA4/m/s5D_8wCJAAAJ
- WAT extractor: Document title bug ia-web-commons#36
...and was fixed for CC-MAIN-2024-42
by commoncrawl/ia-web-commons#37.
This erratum affects all crawls from CC-MAIN-2013-20
until CC-MAIN-2024-38
.
Erratum:
Incorrect fetch_time metadata
In crawls CC-MAIN-2016-36
to CC-MAIN-2016-50
, and CC-MAIN-2018-34
to CC-MAIN-2019-47
the fetch_time metadata for robots.txt
might be incorrect. The correct times can be found in collinfo.json. See the related issue (commoncrawl/nutch#14) for more information.