Microsoft’s top executive privately described AI scraping as 'theft,' according to new unredacted court filings. The admission comes from The New York Times’ copyright lawsuit against OpenAI and Microsoft, which has been ongoing for three years. The documents detail how the companies allegedly bypassed paywalls, built training datasets via mass scraping, and stripped copyright notices from content.
The unsealed material also shows Microsoft’s Copilot 'answer engine' caused a 93% drop in click-through rates for The New York Times’ domain compared to traditional Bing search. An internal Microsoft presentation from January 2024 described the decline as a 'doom loop' that would 'hurt the performance of our models and the entire web at the same time.'
Microsoft CEO Satya Nadella testified earlier this year that 'anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training.' He also said he would have 'invoked [Microsoft’s right] to require OpenAI to retrain its models' if he had known about the scraping.
OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an 'existential threat' from products like the chatbot, which are 'largely substitutive' and 'will get more and more substitutive as they get better.' OpenAI President Greg Brockman described the models as 'excellent at news.'
A Microsoft document states that there is a 'real risk' that generative AI could 'significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.' The documents reveal that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting.
The filings show how OpenAI and Microsoft acquired the plaintiffs’ content, including scraping it from the Bing Index. OpenAI delivered the entire GPT-3 training dataset to Microsoft, which used it to evaluate how to implement OpenAI’s models within its own commercial products. Microsoft also provided training data to OpenAI through initiatives called Project Taxi and Project Mango.
Source: techcrunch