Common Crawl maintains a free,open repository of web crawl data that can be used by anyone.
Common Crawl is a 501(c)(3) non–profit founded in 2007. We make wholesale extraction, transformation and analysis of open web data accessible to researchers.
We analyzed the contents of 584,107 llms.txt files from the July 2026 crawl. Two thirds are produced by a plugin, half follow the structure the specification defines, and a few files even contain prompt injections.