We May Be Scaling the Wrong Layer of AI

The market sees one kind of efficiency coming. There are two and they don't arrive on the same timeline.

David H. Friedel Jr./ 2026-03-25
Subscribe
AIInfrastructureMarkets

In Cheap Compute Is the New Cheap Ride, I argued that the forces behind falling token prices aren’t aligned, that Western monetization gravity and Chinese subsidy logic are pulling developers into different orbits without them knowing which one they’re in.

But there’s a prior question the market still isn’t asking… not just who is making compute cheaper, but how efficiency actually arrives.

Because embedded in that consensus are two very different assumptions about where efficiency comes from. And those assumptions lead to completely different capital outcomes.

The dominant view is straightforward. We build the infrastructure first. Efficiency emerges from scale. More GPUs. More data centers. More energy. Chips improve, inference gets cheaper, costs come down. It’s a familiar pattern, we’ve seen it before.

But this time, there’s a sequencing problem. The system requires massive upfront capital to build the infrastructure that is supposed to eventually make that infrastructure affordable. Capital is front-loaded. Returns are delayed. Efficiency is gated behind scale.

Which means the entire thesis rests on an assumption most people aren’t examining closely enough…

That efficiency is primarily a function of hardware scaling.

The Other Efficiency

There is another form of efficiency emerging, one that doesn’t require the same buildout.

It’s not coming from bigger clusters. It’s coming from better structure… routing tasks to the right-sized model instead of the largest one. Breaking problems into sequential steps instead of brute-forcing them through a single call. Running inference locally where the task allows it. Eliminating unnecessary calls entirely.

If a 1B parameter model, routed correctly, can perform 90% of a 1T model’s task at 0.1% of the cost, the scaling laws haven’t failed. The economic logic of applying them has changed.

This is already happening at the edges. Anthropic’s own model tiers — Haiku, Sonnet, Opus — are an architecture for structural efficiency. So is every developer who wraps a cheap classifier around an expensive reasoning call, or caches deterministic outputs instead of regenerating them. The efficiency isn’t coming from faster silicon. It’s coming from not needing the silicon in the first place.

And it introduces a very different loop.

These loops don’t just differ in mechanism. They differ in what they imply about capital allocation.

Why This Cycle Is Different

In previous infrastructure cycles, efficiency followed scale. You needed railroads before industrial expansion. You needed the internet before SaaS. You needed cloud before modern software distribution. The buildout was prerequisite. Efficiency was downstream.

This time, part of the efficiency is emerging before the system is fully built.

That changes everything. Because now, efficiency doesn’t just depend on infrastructure, it can reduce the need for it.

In The Requisite Chip1, the argument is that the capex cycle is rotating, from GPUs to CPUs, as agentic workloads demand a different kind of compute. That’s a real signal. But it’s still operating within the hardware-efficiency frame. The rotation assumes the buildout continues; it just needs to be rebalanced.

Swapping GPUs for CPUs is still a hardware answer to what may be a software question.

The structural-efficiency frame asks a different question. Not which silicon, but how much silicon. If better orchestration, smarter routing, and local inference meaningfully reduce centralized compute demand, then portions of the buildout, GPU or CPU, become unnecessary regardless of what chips are in the racks.

The most expensive buildings in technology history may be solving for the wrong bottleneck.

The Fork

If hardware efficiency dominates, demand for centralized compute continues rising, infrastructure buildout remains justified, and pricing eventually finds a floor as supply catches up.

If structural efficiency accelerates first, demand for centralized compute softens, portions of the buildout become stranded, and capital is misallocated, not because the technology was wrong, but because the sequencing was.

This is the fork. And the market is largely positioned around one side of it.

The Obvious Counter

There is a version of this argument where structural efficiency doesn’t reduce compute demand at all.

Jevons Paradox

When a resource becomes cheaper to use per unit, total consumption increases because it becomes viable for use cases that were previously uneconomical. If routing a 1B parameter model correctly can do 90% of what a 1T parameter model does at 0.1% of the cost, you don’t get less compute demand. You get a million new agents running tasks that nobody would have paid for at the old price.

This is real. It will happen.

But the question for capital allocation isn’t whether total compute consumed goes up. It’s where that compute runs. If the Jevons demand materializes at the edge, on local models, small inference, lightweight orchestration, then centralized infrastructure still loses the bet. The data centers don’t go dark. They just don’t fill at the rate the financing assumed. If structural efficiency slows the fill-rate of these centers by even 20%, the financing models for hundred-billion-dollar clusters begin to break.

The stranded asset isn’t an empty building. It’s a full building earning below its cost of capital.

Total compute can go up while centralized infrastructure value goes down.

The Mispricing

The market is treating efficiency as a single variable. It isn’t.

If the assumption that scale is the primary driver of efficiency is even partially wrong, then we are overbuilding centralized capacity, underestimating architectural gains, and mispricing where margins will ultimately settle.

This is not a binary outcome. Both forms of efficiency will exist. But the order they arrive in determines where capital flows, which companies win, and whether today’s infrastructure pricing converges upward or downward.

The question is no longer will AI get cheaper?

It is rather, does efficiency come from scale or from structure?

Because if it’s the latter, even partially, we are not just early.

We may be scaling in the wrong direction.

Footnotes

  1. The Requisite ChipGPU-to-CPU rotation in agentic infrastructurehttps://thesynthesis.ai/journal/the-requisite-chip.html — The Requisite ChipGPU-to-CPU rotation in agentic infrastructurehttps://thesynthesis.ai/journal/the-requisite-chip.html
Back to the Journal