Reddit sues AI company Perplexity, others for industrial-scale scraping comments (apnews.com)

🤖 AI Summary
Reddit has filed a federal lawsuit accusing AI answer-engine Perplexity and three data-service operators (Oxylabs, SerpApi and a domain called AWMProxy) of running an “industrial-scale, unlawful” economy that scrapes millions of Reddit comments for commercial AI training. The complaint alleges defendants bypassed Reddit’s anti-scraping controls, masked identities and locations, and even harvested Reddit content via Google Search results — then sold that data to customers like Perplexity. Reddit asserts claims including unfair competition, unjust enrichment and copyright violations; Perplexity, Oxylabs and SerpApi say they’ll vigorously defend themselves and deny wrongdoing. The case matters because it pivots from suing large AI model makers to targeting the scraping supply chain many trainers rely on. If Reddit prevails, it could reshape how conversational datasets are collected, accelerate licensing deals (Reddit already has paid arrangements with Google and OpenAI), and raise legal and operational risk for companies that buy bulk web data. Technically, expect intensified efforts on both sides: more advanced anti-scraping detection and legal barriers from publishers, and more aggressive identity-masking, proxy networks and third-party aggregators from data purchasers. The suit signals heightened regulatory and litigation scrutiny over how publicly available text is repurposed for ML training and could set important precedents about ownership and permissible use of internet-sourced training data.
Loading comments...
loading comments...