Cheap Compute Is the New Cheap Ride

Token prices are falling. The gravity behind the fall is not what you think.

David H. Friedel Jr./ 2026-02-20
Subscribe
AIMarketsInfrastructure

There was a stretch of time when Uber felt structurally mispriced. Not discounted. Not promotional. Simply cheaper than the system it was replacing. You stopped asking what transportation should cost and began assuming the app had solved it. Taxis didn’t disappear because regulation changed overnight; they faded because expectations did.

In hindsight, what was being purchased wasn’t transportation efficiency. It was a behavioral reset. The subsidy trained the habit first. Economics came later.

What looked like generosity was infrastructure for habit formation.

AI feels similar right now; only the subsidy is more complicated, and the forces behind it are pulling in different directions for different reasons.

Two Deflations, One Floor

Instead of sovereign capital underwriting rides, we have deflationary token curves underwriting compute. Cost per million tokens falls. Models compress. Hardware improves. Throughput increases. On paper, everything gets cheaper. If you isolate the compute line item, the trend looks clean and relentless.

But the floor isn’t being held by a single hand.

Western AI platforms are deflating toward eventual monetization. Chinese providers, DeepSeek and the models following in its wake, are deflating toward something structurally different. They are not on an IPO trajectory. They are not building toward pricing normalization. They may be subsidized indefinitely by strategic national interest, which means their incentive to keep prices low is not a prelude to a toll. It may simply be the strategy itself: market share, ecosystem dependency, geopolitical signaling, or proving capability parity at scale.

That distinction matters enormously for anyone building on top of cheap tokens today.

You are not necessarily inside one gravity well. You may not know which one you’re in.

Unit prices fall while the reasons for the falling diverge.

Deflation Is Real, But It Isn’t the Whole Game

What Western platforms increasingly monetize is everything surrounding the token: orchestration layers, reliability guarantees, identity management, tool calling, rate control, compliance, and workflow standardization. The usable system, not the raw generation.

That distinction allows two realities to coexist: unit prices fall while structural boundaries harden. Compute deflates. Constraints professionalize.

The hardening is not accidental. It is incentive-aligned.

Meanwhile, Chinese model availability puts real pressure on Western providers to maintain cheap tokens longer than their economics would otherwise prefer. The competition is genuine. That competition delays the Uber moment, but it does not eliminate it. It extends the window of cheap assumptions, which may make developers more exposed when the shift eventually comes, not less. The longer the cheap era lasts, the deeper the architectural debt.

IPO Gravity Changes the Atmosphere

Pre-IPO companies optimize for growth, but they also optimize for future monetization credibility. Public markets reward predictability. They value systems that can be priced cleanly, forecast reliably, and governed consistently.

That shift rarely appears as a dramatic policy change. It shows up in tone first. Documentation gets clarified. “Recommended paths” become more formal. Authentication methods acquire qualifiers. Nothing breaks in a way that forces outrage, but the direction becomes visible to anyone watching carefully.

This is the rehearsal phase.

Uber didn’t suddenly make rides expensive. It normalized surge pricing after the behavioral transition was complete. In AI, token prices may never spike meaningfully. They don’t need to. If reliability, orchestration, and production-grade guarantees become the default expectation, monetization migrates without the optics of a price hike.

Deflation in raw inputs can coexist with rising effective costs of reliability.

Developers Don’t Pay in Dollars First

Riders pay per trip. Developers pay in architecture.

When you build around permissive access, loosely defined boundaries, or consumer-grade credentials that “just work,” you are investing time into an assumption. That time becomes sunk cost long before any invoice changes.

The most expensive part of AI won’t be the token. It will be the assumption.

Which is why small documentation shifts can trigger disproportionate reaction. The anxiety isn’t about today’s restriction; it’s about tomorrow’s rigidity.

The invoice rarely tells the whole story. The rewrite does.

As a company approaches liquidity, tolerance for ambiguity declines. Not maliciously. Structurally. And the developer who built assuming the Chinese alternative would keep Western platforms honest may find, someday, that the geopolitical calculus shifted and the floor moved for reasons that had nothing to do with compute economics.

Why Deflation Creates Cover

Falling token prices provide strategic cover regardless of their origin. A Western platform can tighten orchestration and formalize identity while truthfully stating that compute has never been cheaper. A Chinese provider can maintain open access while building ecosystem dependency through tooling, fine-tuning pipelines, and local deployment patterns that quietly become load-bearing.

Neither toll needs to be announced.

The reliable path becomes the monetized path. The sovereign path becomes the dependent path. The rest remains technically available, but increasingly peripheral.

Uber did not make every ride expensive. It made the guaranteed ride expensive.

AI may follow the same pattern, from two directions at once.

The Builders Who See It Early

None of this requires cynicism. Subsidies are strategies. Early-phase generosity is rarely equilibrium, and the current moment is unusual precisely because two different subsidy logics are competing to shape your habits simultaneously.

The mistake is assuming that cheap tokens reflect a stable competitive equilibrium rather than a temporary confluence of divergent incentives.

The builders who navigate this well abstract providers early. They separate compute cost from workflow dependency. They design assuming boundaries will harden over time, from the monetization side or the geopolitical side, not because they distrust a platform, but because they understand incentive gravity.

Token prices are deflating. The reasons are not aligned. The duration is uncertain. The exit conditions are different.

Cheap rides taught people how to move before pricing normalized.

Cheap tokens are teaching developers how to build before the subsidies, wherever they come from, whatever they’re for, run their course.

When the assumptions expire, the cost won’t show up as a token increase. It will show up in rewrites. And you may not know, until then, which gravity well you were actually inside.

Back to the Journal