AI models like those powering ChatGPT, Gemini, and Claude are trained on vast databases of published works, including hundreds of millions of books, online articles, and academic papers. Most authors have contributed to these models without their knowledge or consent, raising legal concerns about the use of copyrighted material. The legal landscape surrounding AI training remains unclear, with courts struggling to apply outdated copyright laws to modern technology. 'It’s very complex and there are a lot of raw feelings about what is happening, both for and against,' said Cathy Gellis, an intellectual property attorney. The legal uncertainty has led to significant financial and reputational risks for AI companies, as seen in recent court rulings that have shaped the industry's trajectory.

Last year, Judge William Alsup ruled that Anthropic’s AI training was lawful, despite ordering the company to pay a $1.5 billion copyright settlement for using works from illegal online shadow libraries. 'Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,' the judge wrote, drawing a parallel between AI training and literary study. While the fine was substantial, Gellis noted that it was negligible for a company projecting $200 billion in annual revenue by 2028. She argued that the ruling was more favorable to AI companies, as it emphasized the difference between using and copying copyrighted works.

Copyright law, last updated in 1976, has not kept pace with the rapid development of AI technology. 'Everybody is very worried right now because the law is all over the place, and it’s because of this question,' said Jason Henderson, a senior attorney. 'They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.' The legal debates often center on fair use, which allows for the use of copyrighted materials without explicit permission if the use is deemed transformative. However, courts are divided on how to apply this standard to AI training, with some rulings favoring companies that do not directly compete with content creators.

Source: techcrunch