← Blog
22

2026-01-22 · Bertrand Gonthier

Your data is the product. Again.

Except now it’s training models, not ads.

Everyone’s arguing about which model is “best.” Cute.

The real arms race is who has the legal right to learn from the world—and who gets sued into bankruptcy for pretending “the internet is public domain.”

The split: two AI economies are forming

1) Licensed AI (enterprise-safe, boring, expensive)

  • Built on contracts, provenance, and “approved” corpora.

  • Wins regulated markets because legal departments prefer sleep over benchmarks.

  • Music is already moving this way: the majors are signing AI licensing deals (e.g., Klay with UMG/Sony/Warner). 

2) Scraped AI (fast, cheap, legally radioactive)

  • “We found it online” as a strategy.

  • Great until litigation turns your training set into a liability schedule.

This isn’t hypothetical anymore. Lawsuits are expanding, including publishers pushing to join actions tied to training data for Gemini. 

The new benchmark is: “Can you ship this without a lawsuit?”

Case in point: Anthropic’s $1.5B settlement tied to claims about pirated books used for training. That number is the market screaming: data provenance matters now. 

Canada isn’t watching from the sidelines

Courts

A group of major Canadian news organizations sued OpenAI in Ontario (filed 2024). Reuters covered the action and its claims (damages + injunction). 

Separately, Ontario courts have affirmed jurisdiction for the case to proceed—meaning “you’re not in Canada” is not a magic shield. 

Policy

The Government of Canada’s consultation on copyright in the age of generative AI makes the fault line explicit: creators want consent/compensation; tech groups want broad text-and-data-mining flexibility. Canada hasn’t picked a clean winner yet—and that uncertainty is a business risk. 

Music is the preview of the endgame

Two signals at once:

  • Platforms are drawing a line: Bandcamp just banned AI-generated music (mostly) to protect trust and creators. 

  • Labels are shifting from “block it” to “license it” (structured deals, official channels, pay-to-play training). 

This is where every industry goes next: a rights layer sitting on top of content.

What founders should do (if they like staying alive)

  • Stop saying “we don’t store data” if you can’t prove it. Align retention, logs, and deletion with reality.

  • Create a Data Provenance Sheet: sources, rights basis, retention, opt-outs, and how you handle takedowns.

  • Segment your product: “clean mode” for enterprise + “experimental mode” for hobbyists—don’t mix them.

  • Assume disclosure obligations are coming (EU has GPAI obligations starting August 2025; regulators are not slowing down). 

The uncomfortable conclusion

The next winners won’t just be the labs with the smartest weights.

They’ll be the ones with the best contracts, best provenance, and lowest legal entropy.

Debate question: If “licensed AI” becomes the premium tier and “scraped AI” becomes the shadow tier… do we get a healthier internet, or just a future where only rich models are allowed to learn?


Reading pack

Publishers seek to join lawsuit against Google over AI training

Bandcamp Announces Ban on AI Music

ElevenLabs made an AI album to plug its music generator

Taylor Swift label UMG inks licensing deal with China's NetEase Cloud Music

EU sticks with timeline for AI rules

First measures of European AI Act regulation take effect

Letters from the studio.

One quiet dispatch a month — new work, applied AI notes, no noise.

Have a workflow to fix?

An AI engineer replies within 24 h.

Talk to a human