Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) non–profit founded in 2007. We make wholesale extraction, transformation and analysis of open web data accessible to researchers.
Common Crawl has joined Project Tapestry, a global initiative led by the AI Alliance to advance open, sovereign AI. We will contribute our expertise in responsible web data, multilingual coverage and culturally informed AI development.