The Data
Overview
CDXJ Index
URL Index
Web Graphs
Latest Crawl
Crawl Stats
Graph Stats
Errata
Resources
Get Started
AI Agent
Blog
Examples
CCBot
Infra Status
Opt-Out Ledger
FAQ
Community
Research Papers
Mailing List Archive
Hugging Face
Discord
Collaborators
About
About
Team
Jobs
Privacy Policy
Terms of Use
Search
AI Agent
Contact Us
Read about the Increase of Common Crawl citations in academic research
Research Papers
Open English–French datasets show neural quality filters amplify benchmark contamination
Nathan Godey, et al.
Gaperon: A Peppered English-French Generative Language Model Suite
The largest validated African text-and-speech dataset to date
Sheriff Issaka, et al.
The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP
How LLMs erase Northeast Indian languages
Badal Nyalang
Stereotyped by Silence: How LLMs Erase Northeast Indian Languages Through Omission and Orthographic Corruption
Evaluating AI agent’s capacity to solve real-world CAPTCHA
Xiangyu Wu, et al.
MirrorCAPTCHA: Wild CAPTCHA, Wild Distribution, Wild Web-based Platform Meet Multimodal LLM Agents
What the crawler keeps, and what it loses
Michael Paris, Hande Celikkanat, Luca Foppiano
Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls
Tracking AI adoption across European firms using their websites
Julio Garbers, Terry Gregory
The Diffusion of Artificial Intelligence Across Firms: Evidence from Europe
A new benchmark for web language identification
Pedro Ortiz Suarez, et al.
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Geolocating and embedding 50M German news articles for semantic analysis
Lukas Kriesch, Sebastian Losacker
A geolocated dataset of German news articles
A study on web crawlers facing inconsistent and poorly-signalled blocking
Mostafa Ansar, Anna Sperotto, Ralph Holz
Web Crawl Refusals: Insights From Common Crawl
Research on Free Expression Online
Jeffrey Knockel, et al.
Banned Books: Analysis of Censorship on Amazon.com
Improved Trade-Offs Between Data Quality and Quantity for Long-Horizon Model Training
Dan Su, et al.
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Web Graph Strategies Against Unreliable News
Peter Carragher, Evan M. Williams, Kathleen M. Carley
Misinformation Resilient Search Rankings with Webgraph-based Interventions
Analyzing the Australian Web with Web Graphs: Harmonic Centrality at the Domain Level
Xian Gong, Paul X. McCarthy, Marian-Andrei Rizoiu, Paolo Boldi
Harmony in the Australian Domain Space
The Dangers of Hijacked Hyperlinks
Kevin Saric, et al.
Hyperlink Hijacking: Exploiting Erroneous URL Links to Phantom Domains
Enhancing Computational Analysis
Zhihong Shao, et al.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Computation and Language
Asier Gutiérrez-Fandiño, et al.
esCorpius: A Massive Spanish Crawling Corpus
The Web as a Graph (Master's Thesis)
Marius Løvold Jørgensen, UiT Norges Arktiske Universitet
BacklinkDB: A Purpose-Built Backlink Database Management System
Internet Censorship
University of Maryland, Nourin, Sadia, et al.
Measuring and Evading Turkmenistan’s Internet Censorship
Internet Security: Phishing Websites
Asadullah Safi, Satwinder Singh
A Systematic Literature Review on Phishing Website Detection Techniques
More on Google Scholar
Curated BibTeX Dataset
Text Link