← All episodes
Special · News ·8:18 ·August 2, 2026

How Models are Trained and Will You Pay for It?

An AI company just paid $1.5 billion to settle a copyright case — not for training on the books, but for pirating them. That one distinction is the whole fight over how AI models are built, and it's landing on every company that uses one.

The Promise

  • Courts keep ruling that training on copyrighted work is fair use — which gives companies that license or deploy models a defensible footing, not open-ended liability.
  • A real licensing economy is forming — News Corp, Reddit, and the data deals reshaping the market — so 'paid-for, provenance-clean' models are becoming something you can actually buy.
PROMISE RISK
Balanced

The Risk

  • The $1.5B bill wasn't for training — it was for how the books were obtained: piracy. Provenance, not capability, is where the legal exposure lives.
  • You inherit a model's data provenance the moment you deploy it, and vendor indemnities quietly fail exactly where you'd need them. That's a board risk, not an engineering footnote.

The $1.5 billion distinction

Last year an AI company paid $1.5 billion to settle a copyright case — the largest such settlement in US history. The surprising part: it wasn’t for training its AI on the books. A court had already ruled that part was probably legal. The bill was for how it got the books — it pirated them. That one distinction is the whole fight over how AI models are trained, and it’s about to land on every company that uses one.

How the models are actually built

Strip away the mystique and a large model is built from scraped data at a scale no one licensed up front: pre-training on the open web, books, news, images, music, and code. The courts are converging on a rule that’s cleaner than the headlines suggest — training on copyrighted work tends to count as fair use; acquiring it through a shadow library does not. Meanwhile the labs are hitting a different wall: they’re running out of fresh data, and pivoting to licensing deals and synthetic data to keep going. News Corp and Reddit signed early, and that licensing economy is now reshaping the market.

The part that reaches your board

Here’s why this isn’t someone else’s problem. When you license or deploy someone else’s model, you inherit its data provenance — including the parts nobody documented. Vendor indemnities look reassuring until you read where they stop, and they tend to stop exactly where the real exposure begins. The promise is a maturing market where clean, licensed models are something you can buy on purpose. The risk is every model already in your stack that was built on data it never paid for.