GPUaaS vs. the Token Factory: When AI Infrastructure Stops Renting GPUs and Starts Manufacturing Intelligence
The first generation of AI clouds monetized scarce accelerators by the GPU-hour. The next generation is moving higher in the stack, selling the output of those accelerators as tokens, requests and agentic workflows. That transition changes who carries utilization risk, how infrastructure is financed, how software creates margin—and eventually why the cost of the megawatt underneath the GPU may become one of the industry's most durable competitive advantages.
A GPU is a machine. A token is the product.
For most of the first AI-infrastructure boom, those two things were economically treated almost as though they were the same. AI companies needed accelerators. Accelerators were scarce. Cloud providers acquired them. Customers rented them. The commercial unit was familiar: dollars per GPU-hour.
That model helped create the modern neocloud. But the AI infrastructure market is beginning to move up the stack. Instead of renting the machine, providers increasingly want to sell what the machine produces—tokens, inference requests, model endpoints and agentic workflows. NVIDIA has started describing this transition explicitly: from Compute-as-a-Service toward Token-as-a-Service.
GPUaaS sells time on the machine. A token factory sells the output of the machine.
First, a terminology problem.
GPUaaS, neocloud, managed inference, AI factory and token factory are often discussed as though they are competing business models. They are not—they describe different layers of the stack. A neocloud is generally a type of company: an AI-native cloud provider built around GPU-intensive workloads. That company can sell several different products. It can rent GPUs. It can operate dedicated inference. It can sell serverless model access. Or it can increasingly monetize the tokens those systems produce.
Four different ways to monetize the same physical infrastructure
Rent the accelerator
Managed GPU inference
Sell model access
Manufacture intelligence
The boundaries are not rigid. One AI cloud can offer all four models simultaneously. The important difference is the commercial unit being sold and where operating risk sits.
The neocloud era began with GPU scarcity.
The original economic opportunity was straightforward: acquire large quantities of hard-to-find accelerators, build specialized infrastructure around them and offer customers faster access and better AI-specific economics than general-purpose clouds. Much of that revenue could be contracted through large capacity commitments. CoreWeave is the clearest public example. Its second-quarter 2026 filing says 98% of revenue came from customer commitments, and describes committed contracts as take-or-pay.
That commercial structure is important. If a customer reserves a large GPU cluster, the provider has transferred a meaningful amount of demand and utilization risk to the customer. The customer pays for contracted capacity. Whether the customer's application extracts every possible token from that capacity is largely a different problem.
GPUaaS can look economically like infrastructure leasing: secure the asset, contract the capacity and get paid for making it available.
A token factory changes the unit of sale.
NVIDIA now describes a different commercial architecture. Instead of GPU × hours × price per GPU-hour, the provider increasingly thinks in terms of tokens × realized price per token. In NVIDIA's May 2026 token-metering framework, the company explicitly describes a progression from infrastructure sold by the GPU-hour toward model APIs, inference endpoints and AI applications metered by tokens, requests or workflows. NVIDIA's language is revealing: when token-level economics become visible across the stack, the AI factory becomes a token factory.
Moving up the AI value stack
The physical GPU does not disappear. The customer increasingly stops buying it directly.
Power
Electrical energy enters the AI facility.
GPU
Accelerators convert energy into computation.
Runtime
Serving software schedules and optimizes inference.
Model
Model architecture converts compute into intelligence.
Tokens
The customer's economic output becomes measurable.
Workflow
Agents increasingly turn tokens into completed work.
Nebius is one of the clearest commercial examples.
Nebius launched Nebius Token Factory in November 2025 as a production inference platform for open-source and custom models. The product sits above the physical GPU layer. Customers can access models through an OpenAI-compatible API, while Nebius handles serving infrastructure, autoscaling, model optimization and production operations. Its current service advertises more than 100 million tokens per minute of throughput, more than 60 supported models and a 99.9% uptime SLA.
The strategic direction became even clearer in March 2026, when NVIDIA agreed to invest $2 billion in Nebius and the companies announced plans to collaborate across AI-factory architecture, inference, fleet management and NVIDIA's newest systems as Nebius targets more than 5 GW of NVIDIA systems by the end of 2030.
A basic GPU cloud wins by sourcing accelerators. A token factory has to extract more billable intelligence from every accelerator it already owns.
CoreWeave shows why the categories are converging.
CoreWeave is useful because the same company now exposes several layers of abstraction above essentially the same GPU-cloud foundation. Its Serverless Inference product is billed by the token—customers do not provision infrastructure in advance. Its Dedicated Inference product gives customers explicit GPU and runtime visibility and preserves per-GPU-hour economics. And customers wanting maximum control can operate inference on CoreWeave Kubernetes Service. That produces a continuum: self-managed infrastructure, managed dedicated capacity and fully managed token consumption.
Who owns the economics?
Select a service layer.
Capacity
GPUaaS
The provider monetizes access to an accelerator. The customer generally decides which model to run and how efficiently to use the rented capacity. Contracted GPU capacity can transfer a significant amount of utilization risk away from the infrastructure provider.
Conceptual comparison only. Commercial structures vary by provider, customer, workload and contract.
Baseten shows that a token factory does not have to own every GPU.
The commercial control point can sit above the physical infrastructure. NVIDIA says Baseten aggregates GPUs from more than 10 cloud providers across dozens of global regions and abstracts them into a unified pool for inference customers. Baseten's orchestration layer decides how that underlying capacity is used. The customer primarily experiences the model endpoint, latency, reliability and price—not necessarily which company's physical server produced the token.
A token factory is not defined by owning every machine. It is defined by controlling the production system that turns machines into a consistent AI output product.
Moving up the stack changes who carries utilization risk.
Suppose a customer rents 1,000 GPUs for three years. If the customer only uses those GPUs efficiently 70% of the time, the cloud provider may still receive the contracted capacity payment—the customer's inefficiency is primarily the customer's economic problem. Now suppose the provider instead sells that customer one trillion tokens. The provider has promised the output, and now the provider has to determine how to manufacture those tokens at the lowest possible cost.
The risk transfer changes with the billing unit
GPUaaS
Customer buys capacity
Token Factory
Customer buys output
That is why inference software suddenly becomes margin.
If the customer pays by GPU-hour, better batching and model serving can help the customer consume fewer GPU-hours—which can eventually place downward pressure on infrastructure demand. If the provider is paid by token, better batching directly lowers the provider's production cost. The same applies to KV-cache management, quantization, speculative decoding, request routing, prefill/decode separation, model placement, autoscaling and scheduling. NVIDIA's own performance data illustrates how quickly this layer can move: the company says B200 inference cost on GPT-OSS-120B fell from approximately $0.11 to $0.02 per million tokens through software improvement over a period when the underlying hardware did not change.
When the provider sells tokens, software optimization stops being merely a performance feature. It becomes manufacturing efficiency.
NVIDIA wants the industry measuring factories by output per megawatt.
A GPU tells you how much hardware is installed. It does not tell you how much usable intelligence the facility produces. NVIDIA increasingly emphasizes tokens per second per megawatt and cost per million tokens as the more useful inference metrics. In one April 2026 SemiAnalysis InferenceX comparison cited by NVIDIA, GB300 NVL72 delivered approximately 2.8 million tokens per second per MW on a DeepSeek-R1 workload at the stated latency target, with TCO of approximately $0.123 per million tokens. These numbers are workload-specific and should not be interpreted as universal token-factory economics. But the direction is important: the question moves from “how many GPUs fit in this building?” toward “how much useful AI output can this megawatt manufacture?”
Power price is not necessarily the biggest cost-per-token variable today.
At today's frontier, hardware generation, model architecture, latency requirements, utilization and inference software can move token cost by multiples. Electricity price usually does not. Using NVIDIA's public 2.8 million tokens/second/MW benchmark simply as an illustrative reference, and assuming a facility PUE of 1.20, an $80/MWh delivered electricity price would contribute roughly $0.0095 per million tokens in electricity cost. At $40/MWh the same calculation falls to approximately $0.0048; at $120/MWh it rises to roughly $0.0143. Those figures are illustrative math, not a forecast of production inference economics.
How power price flows into token cost
Adjust the delivered electricity price. The calculation uses NVIDIA's public 2.8M tokens/sec/MW benchmark as an illustrative throughput reference and assumes 1.20 PUE.
Electricity cost / 1M tokens
$0.0095
Electricity component only. Excludes GPU capex, networking, facility capital, software, labor and financing.
500 MW annual power cost
$315M
Assumes 500 MW facility load at 90% annual load factor.
Annual premium vs $50/MWh
+$118M
Shows why seemingly small $/MWh differences become large absolute dollars at hyperscale.
Illustrative calculation, not an investment or operating forecast. The 2.8M tokens/sec/MW reference comes from NVIDIA's cited DeepSeek-R1 benchmark and will vary materially by model, hardware, latency target, batching and software. The 1.20 PUE and 500 MW / 90% assumptions are illustrative Nistar editorial assumptions.
But that does not mean power is unimportant.
It means there are currently larger sources of variance. If one operator's inference stack produces twice as many useful tokens from the same GPU fleet, a $20 or $30/MWh energy disadvantage may be relatively small. But technology advantages tend to diffuse. The latest GPU becomes broadly available. Inference frameworks mature. Optimizations move into open source. Best practices spread. Model architectures converge. Utilization improves across the industry. When those advantages narrow, the cost stack underneath them becomes more visible.
If everyone can buy the same chip and run the same software, competitive advantage migrates underneath the server rack.
That is when electricity starts behaving like feedstock.
This is where a mature token market could start to resemble other energy-intensive manufacturing industries. Consider aluminum. Two smelters can operate similar technology and produce a relatively fungible commodity, yet one plant can remain structurally advantaged for decades because it has access to inexpensive hydroelectric power. The electricity is not merely an overhead line item—it is a fundamental production input. A sufficiently commoditized token market could eventually exhibit some of the same economics: electricity enters, compute transforms it, tokens emerge. The analogy is imperfect—AI output remains far more differentiated than aluminum—but the direction is useful.
Where competitive advantage could migrate over time
Conceptual illustration only. Bar lengths do not represent measured market shares, weights or Nistar underwriting scores.
At hyperscale, small power differences become enormous dollars.
A 500 MW facility operating at a 90% load factor consumes approximately 3.94 million MWh per year. Every $10/MWh difference in delivered electricity price therefore changes annual energy cost by roughly $39 million. A persistent $40/MWh advantage is worth roughly $158 million per year at that scale; a $60/MWh advantage is approximately $237 million per year. Those values do not mean power dominates current inference TCO. They mean that once the rest of the production stack converges, the remaining structural inputs become financially significant very quickly.
The relevant number is not simply wholesale $/MWh.
A token factory cannot run on a headline energy price. What matters is the all-in cost of firm electricity delivered to compute. That can include generation, transmission, distribution, capacity, demand charges, taxes, losses, firming, backup generation, curtailment economics and facility efficiency. A nominal $35/MWh generation source can therefore produce worse token economics than a genuinely firm $60/MWh delivered solution.
In a token factory, cheap electricity matters. Reliable, fully delivered and efficiently converted electricity matters more.
Cost of capital may be the other durable differentiator.
Electricity is not the only structural input unlikely to fully commoditize. AI factories require extraordinary amounts of capital. Accelerators depreciate quickly, while electrical and cooling infrastructure lasts much longer, so the provider has to finance assets with very different economic lives. Two operators running identical hardware on identical electricity can still have materially different token economics if one finances the infrastructure at substantially lower cost. That means the mature competitive stack may increasingly be cheap energy + cheap capital + excellent operations.
This creates an important financing tension.
Moving from GPU-hours to tokens may increase the provider's upside. It can also reduce cash-flow predictability. A long-term GPU take-or-pay contract can produce relatively visible capacity revenue; pure serverless token demand can be much more variable. Customers can change models, change API providers, bring inference in-house, reduce consumption or migrate to a cheaper open model. CoreWeave itself warns investors that a shift away from committed take-or-pay contracts toward pay-as-you-go consumption could make cash flows harder to forecast and affect margins.
More value capture can mean less infrastructure-style certainty
Capacity certainty
Multi-year take-or-pay capacity can produce relatively predictable cash flow and may align better with asset-level financing.
Higher operating leverage
The provider can capture software and utilization upside but carries more demand, pricing and utilization risk.
Reserved output
Token pricing can be combined with reserved capacity, minimum terms and enterprise SLAs to preserve some contractual visibility.
The market is already building that hybrid.
Together AI's Provisioned Throughput is a good example. Launched in July 2026, the product combines reserved inference capacity, token-based pricing, an uptime SLA and a minimum contract term. In other words: the customer experiences token economics, while the provider retains more predictable capacity economics underneath. That may become an increasingly common structure.
NVIDIA has every incentive to push the market toward tokens.
Consider what happens when a new GPU generation doubles useful inference throughput. Under pure GPU-hour pricing, the customer may simply need fewer GPU-hours to perform the same amount of work, so some of the efficiency gain can show up as pricing pressure. Under token pricing, the same hardware footprint can potentially produce twice as much billable output, and the provider can capture more of the performance gain. NVIDIA illustrates this directly in its own Token-as-a-Service material: using simplified H100 assumptions, it compares roughly $18,400 of annual GPU-hour revenue with approximately $157,680 of token-metered revenue per GPU under its hypothetical assumptions. Those figures are explicitly illustrative—they assume specific utilization, token throughput and $/million-token pricing—and are useful for understanding the commercial mechanism, not forecasting real-world returns.
NVIDIA does not merely want customers asking how many GPUs they can install. It wants operators asking how much revenue each megawatt of NVIDIA infrastructure can manufacture.
The next hardware generation reinforces that logic.
NVIDIA's latest messaging around Vera Rubin is increasingly framed around performance per watt. On one August 2026 AgentX benchmark using DeepSeek V4-Pro, NVIDIA reports Vera Rubin NVL72 at up to 30 times the throughput per MW of GB300 NVL72 at the cited high-interactivity point. That is a highly specific benchmark, not a universal 30× improvement across workloads. But it illustrates where AI-infrastructure competition is heading: the scarce resource is increasingly not simply the GPU—it is the megawatt, and the question becomes how much commercially useful intelligence can be produced inside that megawatt envelope.
Eventually, the newest chip stops being a permanent moat.
Every operator wants the newest hardware. Every serious inference platform wants the best software. Open-source runtimes improve, optimization techniques diffuse, fleet-management systems mature and models get more efficient. The leaders will continue innovating, but industry-wide capability tends to converge around widely available tools. If that happens, the things competitors cannot easily copy become increasingly important: a 20-year advantaged power position, a low-cost capital base, a highly efficient physical campus, a location capable of supporting enormous scale. Those advantages sit below the software stack.
The first AI factories will compete on GPUs. The mature AI factories may compete on electricity.
Not every token needs to be produced next to the user.
Geography will still matter. Interactive consumer applications may require very low latency, regulated workloads may require local data residency and enterprise customers may demand specific regional availability. But a large class of future inference workloads may be much more geographically flexible: scientific agents, coding agents, batch reasoning, synthetic-data production, model evaluation, engineering optimization and long-running autonomous workflows. Those workloads may increasingly follow inexpensive, reliable energy rather than expensive metropolitan real estate. That could create two broad forms of inference infrastructure: latency-oriented intelligence close to the customer, and industrial-scale token production where cost and throughput dominate.
That is where AI infrastructure and energy development converge.
If the AI cloud remains primarily a GPU-rental business, energy is an important operating input. If the AI cloud becomes a manufacturing business whose product is intelligence, energy becomes something closer to feedstock. That changes how the physical campus should be understood. Land, generation, transmission, substations, cooling, firming and capital structure stop being merely the infrastructure underneath the technology company. They become part of the technology company's unit economics.
Bottom line
GPUaaS is not disappearing. Customers will continue to want raw accelerators, training will continue to require large dedicated clusters, sophisticated customers will continue to want full control over their inference stacks, and long-term GPU capacity contracts can provide valuable revenue certainty. But the market is clearly adding another layer. The leading AI clouds increasingly want to monetize tokens, models, agents and completed workflows—moving from renting the machine to owning more of the production process. The reward is greater potential value capture; the price is greater operating, demand, pricing and utilization risk.
Today, the biggest differences in cost per token may come from chips, models and software. Tomorrow, those differences may narrow. When every serious operator can buy the same accelerator and run a similarly optimized stack, competitive advantage migrates toward the inputs that are hardest to replicate: energy, capital and the physical infrastructure capable of converting both into compute at scale.
The first AI infrastructure boom was about securing GPUs. The next may be about how efficiently those GPUs can convert electricity into intelligence—and who captures the margin between the two.
Verified sources
Source record reviewed through September 25, 2026. Benchmark results are workload-specific and should not be interpreted as universal production economics.
May 21, 2026. Primary source for NVIDIA's Compute-as-a-Service versus Token-as-a-Service framework, token-factory economics, billing by tokens/requests/workflows and the illustrative GPU-hour versus token-revenue examples.
Primary SEC source for CoreWeave's committed-contract model, take-or-pay structure, 98% committed-contract revenue contribution in Q2 2026 and its discussion of risks associated with movement toward consumption-based pricing.
November 5, 2025. Primary source for the launch of Nebius Token Factory as a managed production-inference platform for open and custom models.
Current product source for 100M+ tokens per minute, production inference, supported model portfolio and 99.9% uptime SLA.
March 11, 2026. Primary source for NVIDIA's $2 billion investment, collaboration across AI-factory design and inference, and Nebius's target of more than 5 GW of NVIDIA systems by 2030.
Primary product source documenting pay-per-token billing, with no GPU-hour billing and no infrastructure provisioning required ahead of demand.
Primary source describing single-tenant inference, explicit GPU choice, runtime flexibility and per-GPU-hour cost visibility.
Source for Baseten's aggregation of GPU capacity from more than ten cloud providers across dozens of regions and its software layer abstracting those resources into a unified inference platform.
Source for NVIDIA's reported B200 cost-per-token improvement on GPT-OSS-120B and the role of TensorRT-LLM software optimization in improving throughput without changing the underlying hardware.
Source for the cited SemiAnalysis InferenceX comparison showing GB300 NVL72 at approximately 2.8 million tokens/sec/MW and $0.123 per million tokens on the specified DeepSeek-R1 inference workload.
July 8, 2026. Example of a hybrid inference structure combining reserved capacity, token-based pricing, an uptime SLA and a minimum contractual term.
August 24, 2026. Source for NVIDIA's workload-specific AgentX performance claims, including up to 30× greater throughput per MW for Vera Rubin NVL72 versus GB300 NVL72 at the cited DeepSeek V4-Pro interactivity point.
Jay Sivam
Expert insights from the Nistar team on energy infrastructure and hyperscale development.