Sony, EMI, and Warner Chappell filed a copyright lawsuit against Anthropic on Friday, alleging that the same torrenting operation that produced a $1.5 billion settlement with book authors also swept up “thousands upon thousands” of the publishers’ copyrighted song lyrics and sheet music. The complaint names both Anthropic co-founder Benjamin Mann and CEO Dario Amodei individually as defendants.

The stakes go beyond one company’s data pipeline. Every major AI lab has trained on scraped web text of uncertain provenance, and Anthropic’s book settlement had already established that fair use protects the training itself even when the underlying copies were pirated. If music publishers can show market substitution instead, a claim authors could not sustain, the fair use shield that survived the book case starts to look narrower than the industry assumed.

According to the complaint, Anthropic’s “mass campaign of illegal torrenting” began in July 2021, when Mann allegedly used BitTorrent to personally download and upload millions of pirated books from Library Genesis, a piracy site later shut down by the FBI. The lawsuit alleges Amodei approved the torrenting. When a replacement mirror called the Pirate Library Mirror surfaced after Z-Library’s shutdown, The complaint says Mann told staff to torrent that one too, telling colleagues the mirror had appeared “just in time.” An Anthropic staffer allegedly replied, “zlibrary my beloved.”

Publishers cite that exchange, resurfaced from filings in the settled book authors’ case, as evidence that piracy was normalized inside the company rather than an isolated lapse. Their claim rests on bibliographic metadata pulled from the LibGen and PiLiMi catalogs, which they say points to Anthropic having torrented hundreds of titles or more that carry lyrics and sheet music the publishers control.

Anthropic denies training its commercial Claude models on any of the pirated material. A company spokesperson told Ars Technica that this is “the third lawsuit from the same lawyers, recycling allegations from cases already before the courts,” and that training generative AI models is “a transformative fair use” under the ruling in the book authors’ case, Bartz. Anthropic said it will defend itself.

The publishers argue that denial rests on a narrow definition of training. Their complaint alleges Anthropic used a non-commercial model trained on LibGen and PiLiMi text to generate synthetic data and reinforcement feedback which they say went on to shape a shipped Claude model, a pretraining pathway that would let pirated text influence a shipped product without ever appearing in its own training run. They also point to testimony from the book authors’ case describing Anthropic’s continued use of LibGen text to check whether model outputs matched source passages too closely, which they frame as evidence the company still relies on the pirated corpus operationally.

Unlike the book authors, who could not prove Claude’s outputs substituted for their work in the market, the music publishers say AI-generated songs are already charting and displacing songwriters directly. They cite a Copyright Office observation that AI outputs competing in the market for a type of work, including song lyrics, can dilute royalty pools even without reproducing a specific copyrighted work. The publishers are asking the court to order Anthropic to produce a full accounting of its training data and methods.

For any lab relying on the Bartz fair use precedent to justify past scraping, this case is the first real test of whether that shield holds when a plaintiff can show direct market substitution instead of abstract training use, and the discovery fight over Anthropic’s pretraining and RLHF pipeline is worth tracking regardless of the outcome.

Ars Technica’s Ashley Belanger reported this story on August 31, 2026.