01 / The physical stack
A token has a supply chain.
The production of a token depends on a model, inference software, a working cluster, servers, networking, cooling, power, high-bandwidth memory, advanced packaging, and silicon. Investors, lenders, and customers fund nearly every step.
This piece focuses on the neocloud, near the middle of the chain. A neocloud commits money before the machine is productive, installs GPUs into clusters, operates the software layer, and sells the resulting capacity. Its immediate product is a GPU-hour; the token appears farther downstream, after a customer chooses the model, workload, latency target, and batching policy.
Operators, lenders, customers, and vendors build a financial chain beside the physical one. Operators use equity, equipment loans, customer deposits, vendor credit, and contract-backed facilities to fund different parts of the build. The customer contract does two jobs. First, it is revenue. Second, it can support leverage if lenders trust the customer and the contract’s acceptance, performance, and payment terms.
One installed GPU has 8,760 theoretical hours a year. Maintenance, failures, commissioning delays, and outages reduce that figure to available hours. Demand determines which available hours are paid; scheduling and workload design determine which paid hours produce useful work; model architecture and inference software determine how many tokens emerge from each productive second.
That chain creates several ways to measure utilization. An operator can measure installed, available, contracted, billable, productive, or token-producing hours, and each denominator answers a different question. A take-or-pay contract can make the customer owe for an hour without running a job, while storage or network bottlenecks can reduce throughput on a technically available cluster.
02 / The economics of a token
The levers, and the leverage, behind token margins.
Asset value
Economic life is the period over which the hardware loses productive value. Residual value is its estimated value at the end of that period; the model depreciates the difference on a straight-line basis.
Financing8
Advance rate is debt as a percentage of eligible asset cost. The all-in interest rate includes the base rate and lender spread. Loan term is the time to maturity; this model assumes full amortization over that term.
Operating assumptions² ⁶ +
An IT kW-month is one kilowatt of computing load reserved for one month.
- Price path. The default case holds realized price at $4.00 for all four years; it does not model renewal or market repricing. That is a strong simplification. H100 pricing shows the pattern older hardware can follow, although it has not been a one-way decline: Silicon Data reports a roughly $3.33 per-hour median in the second half of 2025 and a $2.53 benchmark in August 2026, about 24% lower. Over part of the same period, SemiAnalysis's transaction-validated one-year H100 contract index rose from $1.70 in October 2025 to $2.35 in March 2026 as capacity tightened. Silicon Data ↗ Latest index ↗ SemiAnalysis ↗ These benchmarks cover different contract terms and market segments; they show both long-run price pressure and shorter periods of scarcity. B200 history is still too short to support a credible four-year price curve. Posted prices may not be realizable for a given region, cluster size, term, fabric, or service level.
- Idle power. Electricity is charged only on sold hours, so the model assigns no power cost to idle capacity. Real idle draw is meaningful and would increase cost.
- Power usage effectiveness. PUE is total facility power divided by the power used by IT equipment. A PUE of 1.2 means the facility uses 1.2 units of electricity for every unit delivered to servers: one unit for computing and 0.2 for cooling and other building systems. Lower is more efficient.
- First-year financing. The financing-inclusive tile uses interest on the opening debt balance. Under the modeled amortizing loan, year one has the highest interest expense and the lowest after-interest margin; both improve in later years if the other assumptions hold. Principal repayment and financing fees are excluded. Debt service assumes equal annual principal-and-interest payments over the selected loan term.
- Cash coverage. CFADS, or cash flow available for debt service, is revenue less modeled operating expense and is shown before taxes. DSCR, or debt service coverage ratio, is CFADS divided by annual principal and interest.
- System and throughput scope. Allocated IT load includes the GPU's share of the whole node, not only GPU thermal design power. NVIDIA rates an eight-GPU DGX B200 at approximately 14.3 kW maximum, or about 1.79 kW per GPU. The model defaults to 1.5 kW and allows up to 2.0 kW. The 2,000 tok/s default is also an assumption rather than a chip rating: published B200 results vary materially with the model, sequence length, latency target, batching, and concurrency. NVIDIA DGX B200 ↗ InferenceX ↗ Dell ↗ The model excludes taxes, deployment delay, working capital, and corporate overhead.
- Equity IRR. The calculation starts with capex less debt as the equity outflow, distributes annual cash flow available for debt service after scheduled principal and interest, and adds residual-value sale proceeds at the end of the economic life after repaying any debt still outstanding. The default three-year fully amortizing loan is repaid before the end of the four-year economic life. It assumes annual distributions and flat operating inputs, and excludes taxes, financing fees, working capital, and corporate overhead.
- Customer prepayments. Large, multi-year capacity contracts may include deposits, reservation fees, or prepaid credits. This customer funding can reduce the debt and equity an operator needs before deployment, but it is not additional revenue. The operator earns the prepaid amount by delivering future service and may have to return it if delivery, acceptance, or performance conditions are not met. Lenders also care whether the operator can spend the cash immediately or must hold it in a controlled account.
From GPU-hour to token
How a GPU-hour becomes a token.
A token inherits the economics of the GPU-hour behind it. The machine is only the first cost: every sold hour must also carry the fixed expense of idle capacity and the electricity used while work is running.
The model uses a $4.00 realized price. Under this illustrative cost structure, small changes in price, capex, uptime, or residual value still decide the result. The conversion from GPU-hours into tokens adds one more variable: effective throughput.
realized GPU price / (effective tokens per second × 3,600) × 1,000,000
asset economic cost before financing / (effective tokens per second × 3,600) × 1,000,000
At $4.00 per GPU-hour and 2,000 effective tokens per second, the buyer pays about $0.56 per million tokens for the GPU time. Before financing, the operator’s economic cost is about $0.41 per million. Watch the gap rather than either figure: closing it takes throughput gains, a lower asset basis, higher utilization, or a different price.
For context, $0.56 is raw modeled GPU time, not an API price. OpenAI’s standard short-context flagship prices currently range from $0.20 to $4.00 per million input tokens and $1.20 to $20.00 per million output tokens. OpenAI API pricing ↗ Anthropic lists $1 to $10 for input and $5 to $50 for output across its current Claude models. Claude API pricing ↗ Those prices also cover the model, serving software, reliability, and provider margin.
The unit economics may be clean, but every unit is messy.
Demand splits into two profiles. Training arrives in large, scheduled blocks and can support long commitments, but the customer list is concentrated. Inference is more fragmented and can grow into a steadier base load. Reasoning models use longer outputs and more test-time compute; agentic workloads add repeated model calls, tool use, and asynchronous jobs. Those patterns may raise aggregate demand while making the mix of latency-sensitive and deferrable work more important to fleet economics.
Better software will produce more tokens from each hour, but suppliers do not have to pass the benefit straight through to token buyers. When power, HBM, packaging, or a particular cluster configuration remains scarce, the owner of that constraint can retain part of the productivity gain. Who keeps the margin is an empirical question, and I would not assume the answer is the token buyer.
The unit economics desk prices an average hour. At the end of the day, the operator never sells an average hour: it sells a particular chip, in a particular cluster, to a particular customer, under a particular contract, for a particular workload.
03 / Markets, contracts, and capital
The funding of a token.
An unsold GPU-hour is perishable and fleeting, but the operator would still owe debt service on the machine behind it irregardless of whether the hour is sold. Market liquidity, contract structure, and residual value are parts of the same funding problem.
What, exactly, is the standardized commodity?
A B200-hour and an H100-hour do not deliver the same throughput, and two clusters using the same chip can still produce different amounts of useful work. Inference software, interconnect, batching, availability, and service guarantees all matter. A useful index needs a performance specification and a settlement method.
Move demand
Batch inference, evaluations, rendering, and checkpointable training can absorb off-peak hours.
Move capacity
Brokers, spot markets, sell-back programs, and secondary assignments can put an unused reservation in new hands.
Move price risk
Reservations, forwards, swaps, insurance, and futures can fix price before the delivery window.
Compute commoditization is early, but there are a number of exciting businesses shipping the pieces: broader discovery, normalized contracts, secondary capacity, performance benchmarks, financial hedges, etc. If those layers deepen, operators place spare capacity faster, lenders underwrite firmer recovery values, and merchant fleets pay a smaller financing premium.
Andromeda ↗, Shadeform ↗, Prime Intellect ↗, Vast.ai ↗, io.net ↗, and Akash ↗ attack different parts of discovery, aggregation, deployment, and orchestration. Compute Exchange ↗ and Forward Compute ↗ are working on forwards, swaps, credit cover, outage protection, and residual-value products. Silicon Data ↗ and Ornn ↗ are building benchmark prices, indices, market data, and forward curves. Internet Backyard ↗ is building metering, billing, contracting, and financing rails.
Regulated derivatives would add another layer. On August 11, 2026, CME Group and Silicon Data announced plans to launch compute futures on October 5, subject to regulatory review. The announcement described a planned launch, not a completed one. CME announcement ↗
How the asset gets funded
GPU orders, customer contracts, site delivery, and debt often have large timing gaps. Customer deposits may arrive before the machine is accepted; lenders can use equipment facilities to fund servers as they are delivered; and operators can use corporate capital to bridge delays or support assets that have not yet found a contract.
01Customer capitalDeposits · prepayments · reservations⌄
Customers are often asked to write large checks before receiving service and to reserve capacity for several years. CoreWeave reported weighted-average prepayments of 15–25% of total contract value across active contracts at year-end 2025, while its committed contracts had a weighted-average duration of about five years. CoreWeave 2025 10-K, p. 31 ↗ Contract duration, p. 5 ↗ The financing value depends on refund rights, acceptance tests, termination clauses, and whether unused but available capacity is still payable.
02GPU-backed asset financeSPVs · equipment loans · leases · vendor credit⌄
GPUs, customer contracts, receivables, and controlled cash accounts can sit inside a special-purpose vehicle and support a loan, equipment lease, or vendor facility. These structures are often limited- or non-recourse to the corporate parent after agreed conditions are met, although completion support, equity cures, repurchase promises, and bad-act carve-outs can bring exposure back. CoreWeave’s March 2026 $8.5 billion asset-level facility, for example, had a parent guarantee limited to specified bad acts. CoreWeave March 2026 8-K ↗ Residual value sits at the heart of the advance rate and amortization schedule: if the contract ends before the asset does, repayment depends on redeployment or resale.
03Corporate capitalRevolvers · term loans · notes⌄
Operators use revolvers to bridge deposits and working capital, delayed-draw facilities to fund servers as they arrive, and notes to support a broader expansion plan. The flexibility comes with cross-exposure among deployments and covenants written against the company rather than one ring-fenced asset pool.
04Common equityVenture · growth · strategic · public⌄
VC, growth-equity, strategic, and public-market investors fund the corporate layer and absorb risks that asset lenders will not. They pay for software development, deposits, commissioning gaps, uncontracted capacity, and residual-value exposure. It has no scheduled repayment, but it is usually the most expensive capital in the stack.
05Contract-backed financingCommitted revenue · assigned proceeds · controlled accounts⌄
A lender can size debt against the remaining payments under a long-term customer contract, with customer proceeds flowing through a controlled account as the loan amortizes. The lender still underwrites the operator because those payments depend on delivery, acceptance, service levels, and termination rights. This is closer to contract monetization than ordinary factoring, which usually finances invoices after the operator has delivered the service. Lambda’s August 12, 2026 announcement gives the rough shape: a $926 million first-lien term loan B, issued at 99.5, priced at SOFR plus 300 basis points, rated Baa2, fully amortizing to December 31, 2030, and secured by GPU infrastructure and cash flows dedicated to an unnamed investment-grade offtaker. Lambda announcement ↗
“Investment-grade offtaker” omits most of what a lender needs to know. A lender needs the customer’s identity, concentration, termination and acceptance rights, minimum payments, service-level remedies, parent guarantees, and the extent to which the deployment can be reassigned. Vendor relationships belong in the same analysis. As an example, CoreWeave says all of its GPUs are currently NVIDIA systems because of customer contract requirements. CoreWeave 2025 10-K, p. 15 ↗ In January 2026, NVIDIA invested $2 billion in CoreWeave Class A stock. Supplier concentration, equity stakes, purchase commitments, and vendor-side backstops can often make the economic loop more circular than a simple customer-versus-supplier diagram suggests. CoreWeave 2025 10-K, p. 131 ↗
Controls when hardware cost appears in reported earnings.
Controls the timing of cash tax deductions.
Measures the loss of earning power and resale value.
The operating analogy: asset-leasing businesses such as United Rentals
Obviously, a neocloud has a lot of software complexity, but at the end of the day it owns a financed fleet. It rents capacity, manages price and utilization, maintains the equipment, and eventually has to redeploy or sell it. Current rental yield is only part of the return. A bad residual-value assumption can erase several years of attractive operating results.
United Rentals · equipment rental ↗
United Rentals defines fleet productivity through rental rates, time utilization, and mix, then reports how much original equipment cost it recovers when used machines are sold. The same split works for GPUs: price and utilization drive current yield, while resale or redeployment determines how much of the original capex was ultimately recovered.
The comparison has limits. United Rentals has a broader customer base, a more varied fleet, and deeper used-equipment markets. A neocloud also operates the software, power, and networking stack and promises delivered performance.Report the fleet by hardware vintage; show contracted, paid, and productive utilization separately; publish the lease-expiry ladder and renewal repricing; reserve for upgrades and failures; and compare actual sale proceeds with prior residual-value estimates. Without those figures, a reader can see the current rental yield but not whether the original residual-value assumption was sound.
A shorter economic life raises the break-even price; a weaker residual value tightens advance rates; tighter financing increases the need for customer prepayment; larger prepayments concentrate counterparty exposure; and thin secondary liquidity makes redeployment harder, weakening residual value again. This is a beautiful but delicate system.
04 / The enduring advantage
Where the moat may be.
The durable advantage is the system that keeps scarce infrastructure productive and financeable across customers and hardware generations.
A neocloud runs cloud software, develops infrastructure, manages assets, underwrites customer credit, and prices perishable inventory. The next financing frontier may be merchant capacity: a diversified book of shorter contracts that produces stable portfolio cash flow without one hyperscaler standing behind it. Reaching that point will require clean performance data and evidence through a hardware and credit cycle.
A working ranking
Capacity earns nothing until the utility energizes the site and the customer accepts it.
A full, diversified book improves both asset yield and credit quality.
Useful throughput and reliability matter more than the GPU badge.
An operator with a lower cost of capital can compound its operating edge, but the operator still needs demand.
Older or stranded capacity needs a second customer and a credible price.
Questions I still cannot answer
Can a diversified merchant book become as bankable as one hyperscaler?
Who is best equipped to own residual-value risk?..looking like NVIDIA for now
Which specification makes a GPU-hour hedge settle honestly?
Can a transaction index become liquid before the underlying contracts standardize?
As throughput improves, which scarce layer keeps the margin?
View dated sources and assumptions +
- Andromeda: posted prices observed August 21, 2026: A100 80GB $0.82, H100 $1.84, H200 $2.29, B200 $3.15.
- U.S. EIA: 2025 average industrial electricity price of 8.62¢/kWh; model rounded to $0.09.
- NVIDIA DGX B200 guide: eight-GPU system power of approximately 14.3 kW maximum, or about 1.79 kW per GPU.
- SemiAnalysis InferenceX: workload-specific B200 throughput and cost for Llama 3.3 70B.
- Dell B200 inference tests: an eight-GPU B200 node produced 13,576–25,241 tok/s on a short-response Llama 3.3 70B workload across tested concurrency levels.
- United Rentals 2025 results: fleet productivity and used-equipment recovery.
- All market prices are snapshots, not comparable firm quotes. Capex, colocation, other opex, throughput, economic life, residual value, and financing terms are illustrative model assumptions.
Tokens are priced like software and produced by infrastructure. The margin belongs to whoever best converts capital committed into useful work delivered without getting trapped by the hours in between.
If you operate, finance, broker, insure, or build software around GPU capacity, I would love to compare notes. This market is exceptionally opaque at times and I’d love to help understand and navigate. 🙏