The copyright battle surrounding artificial intelligence is expanding once again. Sony Music and Warner Music have filed a lawsuit against Anthropic in a California federal court, accusing the company of using copyrighted music material to train its artificial intelligence models.
According to the complaint, Claude models were trained using copyrighted song lyrics and sheet music without the necessary authorization. The publishers also allege that the system can reproduce protected works in substantially faithful forms.
The training data problem
The case puts one of the most important questions surrounding AI development at the center of the debate: which content can legally be used to train an AI model?
AI companies require enormous amounts of data to develop increasingly capable models. For the music industry, however, lyrics, sheet music and recordings are intellectual property with significant commercial value.
Anthropic maintains that its training practices fall within the boundaries of fair use in the United States and has said it is prepared to defend itself in court. Music publishers argue instead that using copyrighted works without permission constitutes infringement.
Why the lawsuit matters
The case is about more than Anthropic. Legal challenges against generative AI companies have been increasing as authors, publishers and creative industries seek clearer rules governing the use of copyrighted material.
The lawsuit from Sony Music and Warner Music also follows other disputes involving the music industry. This could increase pressure on AI companies to negotiate licensing agreements with rights holders or develop stricter approaches to the datasets used for training.
The output problem
Another sensitive issue concerns what an AI model produces after training. If a system generates text or other material that is too similar to existing copyrighted works, another layer of questions emerges around responsibility and the limits of automated generation.
For the technology industry, the challenge is therefore twofold: gaining access to the data needed to train increasingly powerful models while also preventing those models from reproducing protected material or directly competing with the creators of the original works.
A case that could reshape AI
The outcome could have consequences well beyond the music industry. If courts impose stricter limits on the use of copyrighted material during AI training, technology companies may have to rethink their data strategies and rely more heavily on licensed content.
If companies receive broader protection for their training practices, creative industries may instead have to develop new ways of monetizing the use of their work in the AI era.
One thing is already clear: the next phase of the artificial intelligence race will not be decided solely by model performance, but also by who owns the data used to build those models.