🤖 AI Summary
Reddit filed suit in federal court in New York accusing four companies—SerpApi, Oxylabs, AWMProxy and Perplexity—of illegally scraping Google search results to extract and sell Reddit content used to train AI models. The complaint alleges the scrapers harvested billions of Google queries monthly, bypassed site protections like robots.txt, packaged Reddit posts, and sold them to clients including large AI firms; Reddit says it spent “tens of millions” on anti‑scraping defenses and seeks an injunction, damages, and a ban on use or sale of any scraped Reddit data. As evidence, Reddit says a specially crafted “test post” that was only crawlable via Google rapidly appeared in Perplexity’s search results, and notes a fortyfold spike in Reddit citations in Perplexity output after a cease‑and‑desist.
This case matters because it targets a core supply chain for modern LLMs: large-scale, often unlicensed web scraping of human-generated content. If courts side with Reddit, AI developers may face stronger legal and technical constraints on using web-scraped datasets, increasing demand for licensed feeds and reshaping training pipelines. The suit highlights practical technical tactics (search-result scraping, robots.txt evasion, data “laundering” and resale) and enforcement challenges—many scraper firms are abroad and use workarounds—so outcomes could influence data‑licensing markets, model provenance practices, and industry norms around compensating creators.
Loading comments...
login to comment
loading comments...
no comments yet