🤖 AI Summary
Reddit has sued Perplexity and several web-scraping firms (Oxylabs UAB, AWMProxy, SerpApi) in Manhattan federal court, alleging they illegally circumvented Reddit’s anti-scraping safeguards to harvest user comments and sell that content for training AI. The complaint says Perplexity agreed not to scrape Reddit, received a cease-and-desist in May 2024, yet its citations of Reddit rose “forty-fold” afterward. Reddit accuses Perplexity of routing content through Google search results and third‑party scrapers — effectively bypassing robots.txt and other defenses — and claims that scraped material underpins a company now valued at roughly $20 billion.
The case is significant because it targets the supply chain that feeds large language models: not just model builders but the intermediaries that crawl and repack public web content. If the suit succeeds, it could force clearer rules around scraping, the legal enforceability of robots.txt and digital guardrails, and new licensing or compliance requirements for training data. Technically, the complaint highlights specific tactics — using search-engine indexing as a backdoor and purchasing scraped corpora — and underscores why platforms invest tens of millions in anti‑scraping systems. The outcome could reshape how AI firms source web data, regulate intermediary scrapers, and balance openness with intellectual property and user-consent obligations.
Loading comments...
login to comment
loading comments...
no comments yet