Reddit sues Perplexity for scraping data to train AI system (www.reuters.com)

🤖 AI Summary
Reddit on Oct. 22 filed a federal lawsuit in New York accusing AI startup Perplexity and three third-party firms of unlawfully scraping Reddit’s content to train Perplexity’s “answer engine” — a search/QA system that relies on human-generated text. The complaint alleges the scrapers circumvented Reddit’s data-protection measures to obtain large volumes of posts and comments Perplexity “desperately needs.” Reddit framed the activity as part of an “industrial-scale ‘data laundering’ economy” and joins a growing wave of content-owners pursuing litigation over alleged use of copyrighted material to train generative AI; Reddit is already suing Anthropic in a separate, ongoing case. For the AI/ML community this reinforces legal and ethical risks around assembling training corpora from public web sources. Technically, the dispute centers on whether automated scraping that bypasses access controls constitutes unlawful copying and whether that data can be used to produce commercial models or services. Outcomes could force startups to rely more on licensed datasets, change scraping practices, or adjust model-training pipelines to avoid contested sources — potentially raising costs and altering performance trade-offs tied to “quality human content.” Perplexity responded that it pursues a “principled and responsible” approach to factual AI and defended openness, signaling the clash between innovation, data access, and copyright law will shape future model development.
Loading comments...
loading comments...