Anthropic has reached a $1.5 billion copyright settlement with book authors, following a federal court in San Francisco's approval of the agreement. The company downloaded books from piracy databases LibGen and PiLiMi between 2021 and 2022, leading to the settlement. Of the approximately 482,460 listed works, 91.3 percent were claimed, with each author receiving about $3,000, four times the statutory minimum. Anthropic must destroy the pirated files and retain claims over AI outputs that reproduce original works and future conduct. This settlement is the largest in class action history, according to the court ruling.
The court ruled that training AI on legally obtained books is 'transformative - spectacularly so' and falls under fair use. However, the ruling does not cover AI training itself, only the piracy aspect. Judge Alsup previously stated that mass scraping of internet content without authors' consent remains an open legal question, meaning the fair use debate is far from over. This decision is seen as a milestone for AI labs that trained on web content without website owners' consent, their main source of training data.
Anthropic must destroy the pirated files and retain claims over AI outputs that reproduce original works and future conduct. The settlement does not cover AI training itself, only the piracy aspect. The ruling does not resolve whether mass scraping of internet content without authors' consent counts as legal acquisition, leaving the fair use debate unresolved. The case highlights the ongoing legal challenges faced by AI labs in balancing innovation with copyright laws.
Source: thedecoder