🤖 AI Summary
New revelations in the ongoing copyright lawsuit brought by The New York Times against OpenAI and Microsoft indicate that top executives have referred to their AI scraping practices as “theft” and acknowledged that AI models pose an “existential threat” to traditional publishers and journalists. Unredacted documents reveal that both companies allegedly bypassed paywalls to gather copyrighted content for training datasets, with Microsoft’s internal communications describing the situation as “the largest theft of labor in human history.” This admission raises significant concerns about the legality of using copyrighted material for training AI, which has been a contentious issue within the industry.
These revelations could have substantial implications for the AI/ML community, particularly regarding the legal landscape surrounding copyright and fair use. Despite judges' previous support for AI firms' claims that training with copyrighted content falls under fair use, internal statements from OpenAI and Microsoft suggest their AIs directly compete with original sources, undermining the very market they rely on. Additionally, the substantial volume of scraped news articles—over 91,000 from The NY Times alone—highlights the potential disruption to the journalism workforce. As these discussions unfold, the AI community must confront the balance between innovation and the ethical use of creator content, raising critical questions about the future of AI training practices.
Loading comments...
login to comment
loading comments...
no comments yet