< Back to Errata

Erratum

Missing Language Classification

Originally reported by 
.

Starting with crawl CC-MAIN-2018-39 we added a language classification field (‘content-languages’) to the URL Index (previously called the "Columnar Index"), WAT files, and WARC metadata for all subsequent crawls. The CLD2 classifier was used, and includes up to three languages per document. We use the ISO-639-3 (three-character) language codes.

Affected Crawls
Affected Web Graphs
No items found.