The Company Quietly Funneling Paywalled Articles to AI Developers (www.theatlantic.com)

🤖 AI Summary
The Atlantic’s investigation reveals that the Common Crawl Foundation — a little-known nonprofit that scrapes and shares petabytes of web data — has become a de facto supplier of training material for major AI labs. Its public archives (updated every few weeks with 1–4 billion pages per crawl) have been used by OpenAI, Google, Anthropic, Nvidia, Meta and others to build LLMs (GPT‑3/GPT‑3.5 included). Although Common Crawl’s site claims it only collects “freely available” content and won’t go “behind paywalls,” its scraper doesn’t execute paywall JavaScript, so it routinely captures the full text of paywalled journalism. Publishers have asked for removals, but the foundation’s immutable file format, misleading search interface, unchanged file timestamps since 2016, and partial deletion claims suggest many articles remain in the archive. CCBot is now the most-blocked crawler among the top 1,000 sites, but blocking can’t purge already stored pages. Technically significant implications: Common Crawl isn’t just a raw dump — it helps curate and host derived datasets (c4, FineWeb, DCLM and others), and its archives power tens of millions of dataset downloads, lowering the barrier to LLM training. That accelerates model capabilities while raising copyright, consent, and economic-harm questions for publishers whose work fuels generative systems. Simple mitigations (attribution metadata, verifiable removal processes) could let researchers keep open access without enabling clandestine commercial training, but Common Crawl’s leadership and recent donations from AI firms complicate incentives.
Loading comments...
loading comments...