AI Training Data: 6 Critical Sourcing Rules for 2026

AI Training Data: 6 Critical Sourcing Rules for 2026

Amazon is buying rare books by the pallet, cutting the spines off, running the loose pages through a scanner, and throwing the paper away. That is not vandalism. It is a supply chain for AI training data — and the destruction is the legally load-bearing part.

404 Media hid a tracking device inside a rare book and followed it to an Amazon facility outside Las Vegas known as VGT3, where bulk book deliveries are debound and digitized. Amazon would say only that it purchases books through commercial channels. Dr. Alex Wissner-Gross flagged the story in The Innermost Loop on August 27, 2026, in a single line: Amazon is slicing spines off books for training data.

Amazon is not being reckless. It is following a playbook a federal judge wrote in 2025. If your company owns content, licenses content, buys AI software, or is thinking about building its own model, that playbook is now your problem too.

AI training data
Lawfully purchased books are destroyed as they are scanned. That destruction is what makes the copy defensible.

What the AI Training Data Ruling Means for Business

In June 2025, Judge William Alsup of the Northern District of California decided the fair use question in Bartz v. Anthropic, No. C 24-05417 WHA. He split the conduct in two, and the split is the whole story.

Buying print books, destructively scanning them, and keeping the digital file was fair use. Downloading the same books from pirate libraries was not. The order is blunt about why: the print original was destroyed, and one copy replaced the other. No net new copy existed, nothing was redistributed, and the format change was therefore transformative on that basis alone.

The pirated library got the opposite treatment. Alsup called that kind of piracy inherently and irredeemably infringing, even where the copy is put to a transformative use and discarded immediately. That half of the case was headed for trial on statutory damages. It never got there. On July 20, 2026, Judge Araceli Martinez-Olguin granted final approval to a $1.5 billion class settlement covering 482,460 works — about $3,000 per book, four times the $750 statutory minimum for ordinary infringement.

The industry read that pair of rulings and drew a precise conclusion about AI training data: the receipt is the defense. Buy the copy, destroy the copy, keep the scan. A warehouse full of ruined first editions is not a scandal. It is a compliance architecture.

The Legal Impact: 6 Rules for AI Training Data

1. Lawful acquisition of AI training data is doing all the work

Under 17 U.S.C. 109, the owner of a lawfully made copy may dispose of that particular copy. Alsup layered a narrow format-change holding on top of it: converting a purchased book to a digital file to save space and enable searchability was transformative for that reason alone. Remove the purchase and nothing else in the analysis survives. If your business is assembling AI training data out of books, manuals, film, images, or third-party datasets, the acquisition paper trail is your case file. Build it before you need it.

2. Buying it later does not cure taking it first

Anthropic eventually bought print copies of many titles it had already downloaded. The court was unmoved — later purchase would not absolve the earlier theft, though it might affect the size of statutory damages. Retroactive licensing is mitigation, not a defense. The same logic reaches every other kind of AI training data you did not pay for on the way in: the scraped spreadsheet, the competitor PDF, and the dataset a contractor brought with him from his last job.

3. The safe harbor is narrower than the headline

Alsup expressly refused to bless everything downstream. He denied summary judgment on copies made from the library for purposes other than training, pointing out that the library had no internal access controls and that hundreds of engineers could copy out of it. One in, one out is the rule, and AI training data pipelines rarely stay that tidy. The moment your scan spawns a second copy, gets handed to a vendor, or leaves the building, you are outside the holding and back in ordinary infringement analysis.

4. Copyright stops at the sale. Your contract has to take over.

This is the part content owners keep missing. The court rejected the argument that authors were entitled to capture an emerging licensing market for training use, and it swatted away the theory that a flood of AI-written competing works is a cognizable harm — comparing it to complaining that schools teach children to write well. Translated: if you publish, license, or sell content, Section 107 will not stop a lawful buyer from training on it. Only your agreement will. That is why an AI training clause is now standard, why how you define confidential information matters more than your privacy policy, and why a well-drafted NDA has quietly become an AI document.

5. Training-data provenance is an M&A diligence item

A $1.5 billion liability grew out of a data-acquisition decision that appeared on no balance sheet and in no cap table. If you are buying a company whose product is AI-built or AI-trained, where the corpus came from is a diligence question with a nine-figure precedent behind it. Ask for the acquisition records. Get representations on provenance, a specific indemnity, and an escrow that survives closing. Our M&A team at Howard Law Group now treats training-data provenance the way it treats environmental history on a real property deal — you do not close without it. The same instinct applies on the sell side: read the legal limits on selling company data before you market the asset.

6. The settlement bought peace on inputs, not outputs

Read the release carefully. Class members gave up claims about the past acquisition and copying of their works — the inputs side — through August 25, 2025. The court went out of its way to note that output-based claims are preserved, and so is anything arising from future conduct. Any vendor that tells you the settlement made its AI training data clean is describing a narrow release as if it were a general one. That distinction belongs in your AI vendor contracts, and our earlier breakdown of the business risks created by the AI copyright settlement walks through the rest.

What Howard East Clients Should Do Now

Four things, in order of how quickly they will bite:

  • Inventory the AI training data you have already fed to a model. Internal documents, client deliverables, licensed stock, purchased datasets. If you cannot produce the license or the receipt for an input, that input is a liability, not an asset.
  • Check your outbound contracts for a training right. Vendors and platforms increasingly take one by default. If you promised a client confidentiality and then granted your CRM the right to train on the same records, you have a conflict you did not negotiate.
  • Put provenance into your diligence checklist. Both directions — when you buy, and when you sell. A clean corpus is now a valuation input.
  • If you are the content owner, stop relying on copyright alone. Register what matters, then do the real work in the license terms. Registration preserves remedies; drafting is what actually controls downstream use.

The uncomfortable through-line is that the law has not fallen behind here so much as it has answered clearly, and the answer favors whoever bought the copy. That is a drafting problem, and drafting problems are solvable.

Talk to Howard East

If your business is training a model, licensing content into one, or acquiring a company that already has, we can review your AI training data contracts and your acquisition trail before someone else does. Book a consultation with Howard East for technology licensing and AI contract work.

This article is for informational purposes only and does not constitute legal advice. Laws vary by jurisdiction and change over time; consult a licensed attorney about your specific situation.

Share This on

Table of Contents

 

 

Howard East is a business-first law firm built for companies and owners who need clear answers, decisive action, and results that hold up under pressure. We focus on complex commercial litigation, corporate and transactional work, and administrative matters—handling everything from deal structure and risk allocation to disputes that threaten the business itself. Our approach is practical and direct: we learn the business, identify the leverage points, and execute a strategy designed to protect your position and maximize outcomes. Clients choose Howard East because we combine high-end legal precision with real-world judgment, responsive communication, and an uncompromising commitment to integrity.

Ready to Protect Your Art and Your Money?

Howard East attorneys work with artists, managers, and creatives on holding company formation, brand deals, IP protection, and outside general counsel retainers.

Related Posts