← Blog
01

2026-07-09 · Bertrand Gonthier

The U.S. Is Losing AI Where It Actually Matters

American AI companies are still winning the headlines, the benchmarks, and the research narratives. But they are quietly losing the only metric that compounds: usage.

Twelve months ago, U.S. models controlled roughly 70% of global demand. Today, they’re closer to 30%. That kind of shift doesn’t happen because of marginal competition. It happens when the market decisively chooses a different default.

You can see it clearly on OpenRouter, one of the few places where real production traffic is visible. Not demos, not benchmarks—actual workloads routed by developers who are paying for every token. In 2024, Chinese models were barely present, sitting around 1% usage. By May, they had crossed 60%. By June, they were processing roughly 18 trillion tokens weekly, compared to about 5.5 trillion for American models. That’s not a slow transition. That’s a system flipping.

Inside China, the scale is even more extreme. Daily token consumption has passed 140 trillion, a more than 1,000x increase in just two years. Entire sectors—insurance, manufacturing, government—are deploying open-weight models on domestic infrastructure. This matters more than it seems. These systems don’t depend on U.S. APIs, don’t rely on Western cloud providers, and can’t be turned off by policy decisions made in Washington. This is not just adoption. It is deliberate insulation.

At the same time, the developer ecosystem is shifting underneath the surface. Stanford’s analysis shows Chinese developers now out-download Americans on Hugging Face. That’s not just a vanity metric. Developers determine what gets built, forked, integrated, and standardized. Once that gravity shifts, it tends to stay shifted.

The immediate driver of all this is brutally simple: price. The leading American models are dramatically more expensive than their Chinese counterparts. Claude runs around $4,811 for a standard evaluation suite. OpenAI is roughly $3,357. DeepSeek comes in at $1,071. Zhipu’s GLM is about $544. At the high end, you are looking at a 5x to 9x premium for performance that is, by most accounts, only marginally better.

That gap might have been sustainable when AI was a feature—something you added to a product for differentiation. It breaks down completely when AI becomes infrastructure. The shift from chat to agents changed the economics overnight. An agent doesn’t call a model once; it calls it thousands of times in a single workflow. Suddenly, cost is not a line item. It is the business model.

This is why, according to Andreessen Horowitz’s Martin Casado, roughly 80% of new AI startups building on open-source stacks are choosing Chinese models. Not because they believe they are superior, but because they are economically viable. If your cost structure doesn’t work, your company doesn’t work.

Even the companies at the frontier are acknowledging the gap is thin. Anthropic has publicly stated that U.S. models are only several months ahead of Chinese systems. That’s the entire advantage. Months. And in exchange for that lead, they are charging multiples. The market is responding exactly as you would expect: it is walking away.

What makes this shift more striking is how poorly timed the strategic decisions have been on the U.S. side. Right as distribution becomes the defining factor, OpenAI launches its strongest model, Sol, and limits access to a small set of government-approved partners. At the exact moment when saturation should be the goal, they chose restriction.

Meta made a different mistake but arrived at the same outcome. It had an early lead in open-weight distribution with Llama, effectively setting the standard for accessible models. And then it stalled. Today, its share of routed production traffic has fallen below 1%. In infrastructure markets, you don’t get to pause. Momentum is the whole game.

Meanwhile, Alibaba’s Qwen has become the most-forked open model in the world. Not because it dominates every benchmark, but because it is cheap, accessible, and easy to deploy. That combination—good enough and everywhere—wins more often than “best but constrained.”

This is the core misunderstanding in Silicon Valley right now. Companies are still optimizing for benchmarks, for marginal gains in reasoning, for leaderboard positioning. The market has already moved on. It is optimizing for cost at scale, deployment flexibility, and ecosystem lock-in. Once a model is embedded into tooling, pipelines, and internal systems, it becomes extremely expensive to replace. At that point, the model is no longer a product. It is infrastructure.

And infrastructure markets have very different rules. They do not reward the best technology. They reward the default. The thing that is everywhere, cheap enough, and deeply integrated.

This is where the story stops being just about companies and starts becoming geopolitical. For decades, U.S. dominance in technology rested on control points: semiconductors, operating systems, cloud platforms. The assumption was that controlling the top of the stack ensured leverage across the entire system.

Open-weight AI breaks that model. If a model can be downloaded, run locally, and adapted without permission, then those control points weaken. Export controls matter less. API restrictions matter less. Even hardware constraints matter less over time as systems are optimized for local deployment.

China’s approach has been consistent with its playbook in other industries: undercut on price, scale aggressively, lock in adoption, and then improve quality over time. It is not subtle, but it is effective. It worked in solar, in telecom, in electric vehicles. There is no reason to believe AI would be different.

By contrast, the U.S. approach increasingly resembles a luxury strategy: high margins, controlled access, and an emphasis on premium performance. That works in markets where scarcity creates value. It fails in markets where ubiquity does.

The uncomfortable reality is that the United States still leads in frontier AI capabilities. But that lead is no longer translating into global usage. And usage is the variable that compounds. It drives data, developer ecosystems, integrations, and ultimately, lock-in.

Once that lock-in sets, switching becomes less a technical question and more an economic one. And economic inertia is far harder to reverse.

What’s happening now is not a sudden collapse but a quiet reordering. The center of gravity is shifting away from where the best models are built toward where the most widely used models live. Those are no longer the same place.

And by the time that distinction becomes obvious to everyone, the outcome will already be decided.

Letters from the studio.

One quiet dispatch a month — new work, applied AI notes, no noise.

Have a workflow to fix?

An AI engineer replies within 24 h.

Talk to a human