Cost Per Token Will Define the Winners in AI Infrastructure
GPU counts, megawatts, and capital expenditure have defined the first phase of the AI infrastructure boom. But those are inputs. As inference becomes the dominant recurring workload, the decisive metric becomes the cost to produce a useful token — and that increasingly favors a vertically integrated approach.
For the last several years, the AI infrastructure market has been measured primarily in inputs.
How many GPUs?
How many megawatts?
How much capital expenditure?
How many data centers?
How quickly can the next cluster be deployed?
Those metrics matter during an infrastructure buildout of unprecedented scale.
But they are not ultimately what the customer is buying.
The economic output of AI inference is intelligence delivered through tokens.
As AI moves from an era dominated by model training into one defined increasingly by continuous inference, we believe the metric that will ultimately matter most is much simpler:
What does it cost to produce a useful token?
That shift has significant implications for the entire AI infrastructure industry.
It means the long-term competitive advantage will not necessarily belong to the company with the largest GPU fleet or the lowest advertised GPU-hour rate.
It will belong to platforms capable of combining energy, data center infrastructure, computing hardware and operational efficiency into the lowest sustainable cost per useful token.
And increasingly, that favors a vertically integrated approach to AI infrastructure.
GPU-Hour Is an Input. Tokens Are the Output.
Today's infrastructure market frequently evaluates compute through the price of a GPU-hour.
That makes sense.
GPUs are scarce, expensive assets, and hourly pricing provides a straightforward way to compare access to different systems.
But GPU-hour pricing measures the cost of an input.
It does not necessarily measure the amount of useful work that input produces.
Two systems charging the same amount per GPU-hour can produce dramatically different quantities of tokens because of differences in accelerator architecture, memory bandwidth, networking, model architecture, batching, utilization, software optimization, latency requirements, and overall system design.
Conversely, a newer system can appear more expensive on a per-GPU basis while producing substantially more output.
NVIDIA has increasingly framed inference economics around exactly this distinction. In April 2026, the company described cost per token—not simply compute cost or FLOPS per dollar—as the relevant total-cost-of-ownership metric for AI inference.
The logic is compelling.
If one infrastructure platform produces twice as many useful tokens from the same economic input, its nominal GPU price matters far less.
AI infrastructure is ultimately a production system.
The relevant question becomes the cost of what that system produces.
The Age of Inference Changes the Economics
Training frontier models has driven much of the first phase of the AI infrastructure boom.
Training jobs are enormous, highly visible and extremely compute intensive.
But training occurs periodically.
Inference occurs every time an AI system is used.
Every query. Every coding request. Every generated image. Every enterprise workflow. Every autonomous agent action. Every reasoning sequence. Every customer interaction.
Amazon has publicly made the same distinction, arguing that inference will ultimately represent the overwhelming majority of future AI cost because models are trained periodically but generate inferences continuously across large-scale applications.
That is an important economic transition.
As AI becomes embedded into applications used by billions of people and increasingly by autonomous software agents, the number of inference events can expand dramatically.
At that scale, relatively small differences in unit cost become extremely significant.
A few percentage points of infrastructure efficiency may not appear transformational when evaluating one GPU.
Across billions or trillions of tokens, they can determine the profitability of an entire platform.
Cost Per Token Is Really a Full-Stack Metric
This is where the conversation becomes more interesting.
Cost per token is sometimes treated as a semiconductor-performance metric.
It is much broader than that.
Every token ultimately carries a portion of the cost of the infrastructure required to produce it.
That includes the accelerator. But it also includes power—the electricity consumed by GPUs, CPUs, networking, storage and cooling; the data center—electrical distribution, cooling infrastructure, land, buildings, substations and supporting systems; networking—both the fabric connecting accelerators and the external connectivity required to serve workloads; utilization—an expensive GPU generating tokens is productive, while an expensive GPU sitting idle is not; financing—AI infrastructure requires enormous upfront capital, and the cost of that capital ultimately flows through the economics of the compute produced; and software and orchestration—how efficiently workloads are scheduled and executed can materially change the output generated by exactly the same hardware.
That means the cost of a token is not determined in one layer.
It is the economic result of the entire infrastructure stack.
Power Cost Will Matter More Than Many Expect
During the first phase of the AI infrastructure boom, GPU availability has been such an overwhelming constraint that electricity price has sometimes been treated as secondary.
If a customer urgently needs scarce compute, paying several cents more per kilowatt-hour may seem insignificant relative to the value of obtaining the GPUs.
That logic makes sense in a scarcity market.
It becomes less persuasive as the industry scales.
Consider a continuously operating 100 MW AI facility.
A four-cent-per-kilowatt-hour difference in electricity cost represents approximately $35 million per year of additional energy expense before considering the impact of cooling and other facility loads.
Across a multi-year infrastructure life, that difference compounds rapidly.
And electricity does not simply power the GPUs.
Every inefficiency elsewhere in the facility increases the amount of electricity required to produce the same computing output.
Cooling efficiency matters. Electrical conversion losses matter. Power utilization matters. Rack density matters. Performance per watt matters.
The economic objective therefore becomes more sophisticated than simply securing inexpensive electricity.
It is about maximizing useful compute output per unit of power.
Google's development of its Ironwood TPU illustrates how seriously the largest AI platforms are approaching this issue. Google designed Ironwood specifically for inference and reported approximately twice the performance per watt of its prior TPU generation, explicitly noting that available power had become a constraint on the delivery of AI capacity.
The competition for AI infrastructure is increasingly becoming a competition for what can be produced from each megawatt.
Cheap Power Alone Will Not Win
This does not mean the cheapest electricity automatically creates the cheapest token.
A data center with extremely inexpensive electricity but poor GPU utilization may still have unattractive economics.
So can a facility using outdated accelerators. Or inefficient cooling. Or poorly designed networking. Or hardware mismatched to the workload.
A ten-percent improvement in power pricing can be overwhelmed by a significantly larger difference in accelerator performance or utilization.
The winning formula is therefore not "find the cheapest power."
It is: combine competitive power with the infrastructure and compute architecture capable of turning that power into the greatest amount of useful AI output.
That distinction is critical.
Power becomes part of the cost-per-token equation rather than an isolated real-estate consideration.
Why Vertical Integration Starts to Matter
Historically, the computing stack has contained multiple independent economic layers.
A utility sells electricity. A data center operator purchases the electricity and sells capacity. A cloud provider purchases or leases data center capacity and deploys hardware. A compute provider monetizes the GPUs. An application company purchases the compute and sells an AI product.
Every layer needs its own economics.
Every layer can introduce another margin, constraint, contractual obligation or operational inefficiency.
Vertical integration changes that equation.
A platform controlling the physical stack from power to data center to GPUs has the ability to optimize the system around one economic objective rather than three separate businesses optimizing their individual components.
Power decisions can reflect the actual compute workload. The data center can be engineered around the density and cooling requirements of the hardware. Hardware deployment can align directly with available power. Infrastructure can be phased with customer demand. Capital can be deployed across the stack rather than purchasing each layer at another provider's retail economics.
The opportunity is not simply to capture more margin.
It is to remove friction between the layers.
The Hyperscalers Are Already Moving in This Direction
The industry's largest participants provide an important signal.
Google does not view its TPU independently from its data center.
Its AI Hypercomputer architecture combines custom silicon, networking, cooling and software as an integrated system designed to deliver better performance and economics. Google has described its infrastructure strategy explicitly in terms of optimizing the entire hardware and software stack rather than optimizing one component independently.
Amazon is following a similar path with Trainium.
AWS developed its own AI accelerators because the economics of relying entirely on third-party silicon eventually become strategically important at hyperscale. AWS has stated that Trainium2 provides roughly 30–40% better price-performance than comparable GPU-based EC2 instances.
The point is not that GPUs disappear.
They will remain central to AI infrastructure.
The broader lesson is that the largest operators are increasingly attempting to control more of the infrastructure required to deliver AI because system economics improve when the layers can be optimized together.
Vertical Integration Does Not Mean Every Component Must Be Proprietary
There is an important distinction.
Vertical integration does not require a company to manufacture its own transformer, turbine, GPU and cooling system.
Nor does every AI infrastructure provider need to design custom silicon.
The more important concept is vertical control.
A platform can use best-in-class third-party technology while still controlling the economic architecture around it.
The generator may come from one manufacturer. The GPU from another. The cooling system from another. The networking fabric from another.
What matters is whether those components are being assembled as independent products with stacked economics or as one coordinated infrastructure platform.
That distinction will become increasingly important as AI infrastructure matures.
Utilization May Be as Important as Acquisition Cost
Another reason cost per token is a more useful measure than GPU-hour pricing is utilization.
A GPU is an unusually expensive asset.
Its economics change dramatically depending on how much productive work it performs during its economic life.
A cheaper GPU fleet running at low utilization can be more expensive per token than a more costly system operating efficiently.
This is one of the reasons increasingly sophisticated inference architectures focus not just on raw accelerator performance but on batching, memory utilization, scheduling and networking efficiency.
AWS recently demonstrated that techniques such as speculative decoding can increase token-generation performance significantly for certain workloads without proportionately increasing hardware requirements.
The physical infrastructure provides capacity.
How effectively that capacity is converted into output determines the economics.
The Hardware Cycle Will Keep Compressing Token Costs
There is another dynamic that will accelerate this shift.
AI hardware is improving extremely quickly.
Every new generation is competing on some combination of greater compute density, more memory, higher memory bandwidth, faster interconnect, improved performance per watt, and lower cost per unit of output.
That means the industry cannot assume today's GPU economics remain static.
As hardware improves, token production becomes cheaper.
As token production becomes cheaper, AI applications can consume more tokens.
That increased consumption creates more infrastructure demand.
This is similar to many previous technology cycles: reducing the unit cost of a resource can dramatically expand the market for that resource.
The companies best positioned for that environment are not simply those capable of buying today's hardware.
They are those capable of continuously incorporating better hardware into an infrastructure system whose underlying power and facility economics remain competitive.
Eventually, AI Infrastructure Becomes an Industrial Efficiency Business
This is where we believe the industry is heading.
During a scarcity cycle, access dominates.
Who has the GPUs? Who has the megawatts? Who can deliver first?
Those advantages will continue to matter.
But mature infrastructure markets eventually become increasingly focused on efficiency.
The question shifts from "Who can produce it?" to "Who can produce it most economically?"
Energy markets already operate this way. Manufacturing operates this way. Cloud computing evolved this way.
AI inference is likely to follow a similar path.
When the underlying product is produced at enormous scale, the marginal cost of producing each additional unit becomes strategically important.
For AI, that unit is increasingly the token.
Cost Per Useful Token Will Be the Better Scoreboard
There is one final qualification.
The goal cannot simply be the cheapest possible token.
A token generated too slowly may not meet the application's latency requirement.
A system optimized for maximum throughput may not provide the responsiveness required by an interactive workload.
Different models also generate different levels of intelligence and require different amounts of compute.
That is why the more useful economic concept is cost per useful token.
Useful implies that the token meets the performance, latency, reliability and quality requirements of the workload.
The winning infrastructure platform will therefore optimize multiple variables simultaneously: cost, power, throughput, latency, utilization, and reliability.
That is ultimately a full-stack problem.
Bottom Line
The first phase of the AI infrastructure boom has been defined by scarcity.
Scarce GPUs. Scarce power. Scarce data center capacity. Scarce time.
The next phase will increasingly be defined by economics.
As inference becomes the dominant recurring AI workload and hardware continues to improve, customers will care less about what individual infrastructure components cost and more about what those components collectively allow them to produce.
The ultimate metric will increasingly become the cost of delivering useful intelligence.
Cost per token will become one of the defining economic measurements of AI infrastructure.
Power cost will be part of it. Data center efficiency will be part of it. GPU performance will be part of it. Utilization and software optimization will be part of it.
And the platforms capable of controlling those layers together will have a structural advantage over those purchasing each layer independently.
The future of AI infrastructure is therefore not simply about owning GPUs.
And it is not simply about owning power.
It is about connecting the entire physical stack—from energy, to data center, to compute—and optimizing it around a single unit of economic output.
The long-term winners will not be measured by how many GPUs they own.
They will be measured by how efficiently they turn power into intelligence.
Jay Sivam
Expert insights from the Nistar team on energy infrastructure and hyperscale development.