Microsoft and OpenAI have faced legal scrutiny over scraping news content for AI training, with internal documents revealing the companies' awareness of the harm caused to journalism. News organizations accused the firms of violating copyright laws by using news content to train AI models like ChatGPT and Copilot.
Internal documents show Microsoft's Brent Hecht warned that scraping news for AI training was 'an astonishing theft of unprecedented proportions,' calling it 'the largest theft of labor in human history.' He also argued that the practice 'mocked the idea of fair use,' according to news plaintiffs.
OpenAI's Nick Turley described the use of news content for AI training as an 'existential threat' to publishers, while a Microsoft document warned of a 'doom loop' that could hurt both the models and the web. Data from both firms showed significant drops in click-through rates for news sites, supporting the claim of substitution.
"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain,'" the Microsoft document said. News groups argue that the evidence of substitution could undermine Microsoft and OpenAI's fair use claims.
Microsoft and OpenAI have denied wrongdoing, with a spokesperson stating that the internal documents reflect individual perspectives and not company policy. However, news plaintiffs argue that the documents reveal the companies' awareness of the harm caused by their practices.
Source: arstechnica