On May 11, 2026, technology analyst Ben Thompson published an essay on his influential Stratechery newsletter titled “The Inference Shift,” arguing that the economic center of gravity in artificial intelligence is moving from training — the one-time, compute-intensive process of building a model — to inference, the ongoing work of running that model every time a user asks it a question.
Thompson’s framing matters because Stratechery is widely read by technology executives and investors, and because the training-versus-inference balance directly shapes where the next wave of infrastructure spending — chips, data centers, power, and networks — actually lands.
Executive Summary
The essay’s core contention, as its title signals, is that the AI buildout’s defining workload is changing. Training a frontier model is a bounded project: enormous, but finite, concentrated in a handful of massive facilities run by a handful of well-capitalized labs. Inference is different in kind. It scales with usage — every chatbot session, coding assistant, and AI-powered search query consumes compute — so as AI products find real adoption, serving them becomes a continuous, growing operating cost rather than a one-time capital project.
For infrastructure providers, that distinction is not academic. Training demand rewards maximum-density campuses wherever cheap power and land exist, with latency largely irrelevant. Inference demand rewards something closer to the traditional internet: capacity distributed nearer to users, resilient connectivity, and economics measured in cost per query rather than cost per training run.
Because the full essay sits behind Stratechery’s subscription, this analysis works from the thesis itself — the shift from training to inference economics — rather than from the piece’s specific figures or examples, and examines what that shift would re-rank across the infrastructure landscape.
Two Very Different Kinds of Compute Demand
Training and inference stress infrastructure in almost opposite ways. Training jobs run for weeks or months across thousands of tightly interconnected accelerators, which pushes builders toward gigantic single-site campuses where power is cheap and abundant — remoteness is a feature, not a bug. Inference workloads are short, bursty, and user-facing: a response has to come back in a second or two, which puts a premium on proximity to population centers, redundancy, and network quality.
The economics diverge just as sharply. Training is capital expenditure that a company chooses to make; it can be deferred, right-sized, or cancelled. Inference is tied to revenue-generating usage — if customers are querying your model, you must serve them, and your margins depend on how cheaply you can do it. A market organized around inference is one where efficiency per query, not raw peak capacity, becomes the competitive battleground.
What Gets Re-Ranked in Infrastructure Demand
If Thompson’s thesis holds, several categories of infrastructure move up the priority list. Metro and regional data centers — including colocation capacity near enterprise users — regain relevance after a period in which headlines were dominated by remote gigawatt-scale training campuses. Connectivity providers benefit, because distributed inference multiplies traffic between users, edge sites, and core facilities. Power demand becomes more geographically dispersed and steadier in profile, a different planning problem for utilities than a handful of enormous point loads.
The chip layer re-ranks too. Training has been dominated by the most powerful general-purpose GPUs, where flexibility justifies premium pricing. Inference, being a more predictable and repetitive workload, is friendlier to specialized silicon and to cost-optimized accelerators — which is precisely why cloud providers have invested in custom inference chips and why competition at this layer is more open than in training hardware.
Winners, Losers, and the Margin Question
The clearest beneficiaries of an inference-led market are operators with distributed footprints, strong interconnection, and the ability to sell capacity in smaller, latency-sensitive increments — along with any vendor that reduces cost per query, from silicon designers to cooling and power-efficiency specialists. The more exposed parties are those whose plans assume training demand grows indefinitely on its current trajectory: single-tenant mega-campuses purpose-built for one lab’s training runs carry concentration risk if that lab’s training appetite plateaus while its serving needs move elsewhere.
There is also a margin story embedded in the shift. When inference is the dominant cost, AI application companies face a squeeze between what users pay and what serving costs — which pressures them to negotiate hard with infrastructure suppliers, adopt cheaper hardware, and shrink models where quality allows. Infrastructure revenue may keep growing, but the pricing power within the stack could redistribute.
Reasons for Caution
The thesis has honest counterarguments, and they deserve equal scrutiny. Frontier labs continue to spend heavily on training, and newer techniques that make models “think longer” at answer time blur the line — they raise inference costs, supporting the thesis, but also keep demand for dense, training-class hardware high. It is also possible that both curves rise together, in which case “shift” overstates a rebalancing. And headline-level analysis of a subscription essay cannot verify which evidence Thompson marshals; readers should treat the thesis as a framework to test against disclosed capital-spending and usage data, not as settled fact.
Background
Stratechery, founded by Ben Thompson in 2013, is a subscription publication analyzing the strategy and economics of the technology industry, and it has been one of the more influential independent voices in debates over the AI buildout. The training-versus-inference question it takes up here has become central to that buildout: the industry’s first phase was defined by a race to train ever-larger foundation models, concentrating spending on top-end GPUs and massive single-site campuses.
As AI products have moved from demos to daily tools, attention has turned to the cost of actually serving them at scale. Cloud providers have developed custom inference chips, model developers have released smaller and cheaper model variants, and newer ‘reasoning’ models that consume extra compute per answer have pushed inference costs up further — all of which forms the backdrop against which Thompson’s May 2026 essay lands.
Source: The Inference Shift — Stratechery by Ben Thompson, an analytical essay published May 11, 2026, arguing that AI economics are moving from model training to inference.


