Sony and Warner Sue Anthropic Over Pirated Training Data, Seeking Up to $150,000 per Song
A copyright suit in the Northern District of California names Anthropic's co-founders individually and assigns a separate penalty to stripping metadata from training files.
What happened
Sony Music Publishing and Warner Chappell filed a copyright infringement lawsuit against Anthropic in the U.S. District Court for the Northern District of California, naming co-founders Dario Amodei and Benjamin Mann as individual defendants alongside the company. The suit covers what The Verge reports as "tens of thousands" of copyrighted works — TechCrunch's account says "thousands" — and seeks up to $150,000 in damages per work, plus an additional $25,000 for each instance in which identifiable copyright metadata was stripped.
Context
This is not the first music-industry action against Anthropic. The company already paid $1.5 billion in the Bartz v. Anthropic matter, which The Verge characterises as a settlement and TechCrunch describes as a judge-ordered payment following a ruling that using copyrighted material was legal but acquiring it through piracy was not. Anthropic has also faced suits from Universal Music Group, Concord, ABKCO, BMG, and Round Hill Music. The law firm representing the Sony and Warner plaintiffs in this filing also represents Concord and UMG in a separate case filed in January 2026. Music Business Worldwide broke the story before The Verge and TechCrunch published their reports on 29 August 2026.
How it works
The technical claim concerns data acquisition, not model architecture. The complaint alleges three distinct channels: Benjamin Mann downloaded over five million pirated books via BitTorrent; unnamed Anthropic employees pulled at least two million more from Pirate Library Mirror; and the company scraped lyrics from MusixMatch and LyricFind, two databases that paid the labels to licence their content. Specific tracks named in the filing include "Ain't No Mountain High Enough" (Marvin Gaye & Tammi Terrell), "Livin' On a Prayer" (Bon Jovi), "September" (Earth, Wind & Fire), "Hallelujah" (Leonard Cohen), and "Paper Rings" (Taylor Swift). The $25,000-per-instance metadata penalty treats the removal of attribution data as a separate act of infringement, not merely evidence of one. No source describes the model weights, training framework, or inference stack involved.
Our read
The legal move that matters most here is not the scraping claim — that is the expected one. It is the metadata-stripping penalty. By assigning a separate $25,000 per instance to the act of removing attribution data, the plaintiffs are building a theory that the concealment itself is an independent infringement. That shifts the evidentiary burden: a defendant can no longer argue "we did not know where the file came from" if the metadata was deliberately removed before training. The distinction between "we had it" and "we had it and hid the label" becomes a pricing question.
Naming Amodei and Mann as individual defendants is a deliberate escalation from the Bartz outcome, where the $1.5 billion was a corporate payment. The BitTorrent allegation against Mann specifically — five million books — is written to make individual liability feel personal rather than a compliance failure. Whether the Bartz payment was a negotiated settlement (Verge) or a judicial order (TechCrunch) matters for how a court treats the precedential value of the prior ruling; the sources disagree and the complaint itself is not reproduced in either article.
The scale gap is material. At $150,000 per work, the difference between "thousands" and "tens of thousands" is the distance between roughly $750 million and several billion dollars. The complaint, not either news summary, will settle that figure.
What this changes
For a studio running ComfyUI, local Flux or SDXL video models, or LoRA training on openly licensed image data: nothing changes on Monday. The suit targets text and music training-data acquisition, not video-generation pipelines, and no ComfyUI node graph is implicated.
Where it does touch a small studio's workflow: if you fine-tune a TTS model, a music-conditioned model, or any lyric-aware model on audio or text pulled from commercial sources — streaming services, lyric databases, paid libraries — the acquisition method is now a visible legal exposure, not just an ethical one. The metadata-stripping penalty means re-hosting a track without its ID3 tags does not reset the infringement clock. The concrete action: before the next fine-tune cycle, audit every audio and text file in your training corpus for source provenance. If you cannot point to a licence or a Creative Commons grant, it does not go in the training set.
License
This story is a copyright-infringement lawsuit, not a model or software release. No licence for any Anthropic model, dataset, or weight is named in either source, and no licence question is at issue in the filing. No licence applies.
Key takeaways
- The suit seeks $150,000 per copyrighted work plus $25,000 per stripped-metadata instance, with total exposure reaching several billion dollars at the maximum.
- The Bartz precedent drew a line between lawful use and unlawful acquisition; this suit argues Anthropic crossed the acquisition line on a scale of millions of files.
- Assigning a separate penalty to metadata removal and naming co-founders as individual defendants are legal strategies that extend well beyond the scraping claim itself.
- For studios fine-tuning on scraped audio or lyric text, the method of acquisition — licence versus scrape versus torrent — is now a concrete legal exposure point, regardless of where the model runs.
- The exact number of works in the suit is disputed between sources: "tens of thousands" (The Verge) versus "thousands" (TechCrunch); the complaint itself is not reproduced in either article.
Sources
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 2
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260829T210338Z