A model token is a unit of text a tokenizer emits or consumes. A crypto token is a transferable digital asset. A tokenized deposit is a bank liability recorded on programmable infrastructure. A usage credit is a contractual right to consume a defined service.
All four can be represented in software. None of them are the same instrument. They have different issuers, different obligations, different redemption properties, and different legal status. The fact that a single English word covers all four is an accident of vocabulary, not a discovery about markets.
That accident is now driving strategy. Companies are building compute exchanges on the premise that model tokens are fungible. Investors are pricing routing platforms as if throughput were margin. Vendors are describing usage credits with language borrowed from monetary economics. Each of these moves imports the properties of one category into another, and each will be corrected by contact with an institutional buyer’s procurement team.
This essay argues a narrower and more durable position — the governed-compute thesis. AI compute will not converge on a single universal token. Physical capacity, model processing, enterprise usage rights, and financial settlement will remain distinct. The opportunity that survives the confusion is a governed control layer, one that converts heterogeneous compute into authorized, measurable, auditable work.
Four things called tokens
Start with the category error, because everything downstream depends on getting it right.
The four objects share a name and nothing else. A model token has no issuer and no redemption. It is a measurement artifact — a count that depends on the tokenizer, the model architecture, the input-output mix, context length, caching, batch size, latency target, and quality class. Two workloads that both consume a million tokens can differ in cost by an order of magnitude and in business value by more.
A crypto token has an issuer, a transfer mechanism, and a market price. Its supply schedule is a design choice. Its price responds to speculation as readily as to use.
A tokenized deposit is a claim on a specific bank. The Bank for International Settlements treats two properties as non-negotiable for any monetary arrangement: a common unit of account, and singleness, meaning claims denominated in that unit are redeemable at par with central bank money, with finality. Those properties are what make the instrument money. They are not incidental features a software platform can adopt selectively.
A usage credit is a contractual entitlement. Its issuer owes a service, not a sum. It is not redeemable for cash, and it should not be transferable.
One word, four instruments
Each column is a different economic object with a different issuer, obligation, and redemption property. Only one of them is money.
- Issuer
- None
- Obligation
- None
- Redeem
- Not applicable
- Transfer
- Not applicable
- Issuer
- Protocol
- Obligation
- None fixed
- Redeem
- Market price only
- Transfer
- Freely
- Issuer
- A bank
- Obligation
- Bank liability
- Redeem
- At par, with finality
- Transfer
- Within regulated rails
- Issuer
- The platform
- Obligation
- Service delivery
- Redeem
- No cash redemption
- Transfer
- Non-transferable
The copper column is the only monetary instrument in the set. Borrowing its vocabulary for any of the other three imports obligations the issuer has not actually taken on.
Strategy that treats these as one object produces predictable failures. It markets an entitlement as an investment. It prices a service against a benchmark it cannot control. It promises fungibility across models that are not fungible. And it invites a regulatory question the business was never structured to answer.
Five layers, not one market
The AI compute market is often drawn as a stack with margin migrating upward. That picture is too simple. It is better described as five layers, each selling a different thing, each carrying a different risk, and each measuring in a different unit.
A durable strategy respects these boundaries. Physical units measure capacity. Model units measure processing. Usage credits commercialize governed consumption. Money settles the contract. Every layer can be programmable without becoming the same instrument.
The five layers and their units
Margin does not simply migrate upward. Each layer sells a different object, prices in a different unit, and carries a risk the layers above and below do not.
The copper band is the governed layer. It is the only one whose unit is defined by the customer’s authorization rather than by a supplier’s hardware or a model’s tokenizer.
The distinction is not academic. It determines whether a platform can change providers without renegotiating customer contracts, whether a usage credit can be marketed without securities exposure, and whether a customer experience can be organized around outcomes rather than provider-specific price lists.
What CoreWeave’s quarter actually shows
The clearest public evidence on infrastructure economics is CoreWeave’s first quarter of 2026, and it is routinely misread.
The company reported revenue of $2.078 billion, up approximately 112 percent year over year, against a revenue backlog of $99.4 billion. In the same quarter it reported a GAAP operating loss of $144 million, a negative 7 percent operating margin, adjusted operating income of $21 million at a positive 1 percent margin, $1.147 billion of depreciation and amortization, and $536 million of net interest expense.
Read carelessly, this looks like a price war. It is not. It is capital intensity.
Capital intensity, drawn to scale
CoreWeave Q1 2026, in millions of dollars. Depreciation and financing are two separate pressures on two separate lines — and together they dwarf the margin they are blamed for compressing.
The gap between the (7)% GAAP margin and the +1% adjusted margin is roughly $165 million of add-backs. Anyone citing either figure should say which one, and why. Source: CoreWeave investor relations, Q1 2026.
Depreciation alone equals roughly 55 percent of revenue. That is the cost of buying accelerators fast enough to serve contracted demand, recognized over their useful life. Net interest expense of $536 million sits below operating income and reflects how that hardware was financed. These are two distinct pressures on two distinct lines, and collapsing them into one narrative about pricing produces the wrong conclusion.
Strong demand can coexist with weak operating economics when rapid capacity expansion requires large depreciable assets and substantial financing.
Revenue growth does not eliminate depreciation, financing cost, utilization risk, or customer concentration. A $99.4 billion backlog is a commitment to deliver, and delivering requires more of exactly the capital that compresses the margin.
This is not a prediction that infrastructure providers fail. It is an observation that owning the metal is a capital business with capital-business returns, and that a control layer sitting above it has a different cost structure and a different risk profile. Published GPU pricing reinforces the point: rates vary widely by provider, form factor, region, commitment, spot availability, network fabric, and service level, as market indices such as the SemiAnalysis GPU Cloud Index track in detail. A price index rising from a trough indicates scarcity in a particular contract segment. It does not establish that normalized cost per unit of capability is rising, that all providers are expanding margin, or that the next hardware generation preserves the same rate.
Two prices moving in opposite directions
The most common mistake in AI cost forecasting is picking one direction.
Research from MIT’s FutureTech group, published as The Price of Progress, measures both. Using pricing data from Artificial Analysis and Epoch AI, the authors find that the price of achieving a given level of benchmark performance has fallen roughly five- to tenfold per year. Over the same period, the price of running frontier models has risen between three- and eighteenfold per year, driven by larger models and heavier reasoning demands. Algorithmic efficiency alone contributes roughly a threefold annual improvement.
The scissors
Both statements are true at once, because they answer different questions. Holding capability fixed, price collapses. Chasing the frontier, price climbs. Vertical scale is logarithmic; bands show the reported ranges over three years.
Source: Gundlach, Lynch, Mertens and Thompson, The Price of Progress: Price Performance and the Future of AI (arXiv:2511.23455), using Artificial Analysis and Epoch AI pricing data. Ranges as reported; the chart illustrates direction and spread, not a forecast.
Efficiency reduces the cost of a fixed capability. If you need last year’s performance, you will pay far less for it this year. But demand does not hold capability fixed. It moves toward longer context, tool use, multimodal processing, and agentic loops that consume more tokens along more expensive execution paths. The buyer who benchmarked a workload eighteen months ago and assumed the bill would fall has watched the workload change underneath the assumption.
The commercial consequence is that no vendor should promise a customer that token prices only fall. Nor should any vendor build a business on the premise that they only rise. The real problem is managing a changing mix of workloads against a contracted service envelope. That is a governance problem, not a pricing prediction.
Tokens per watt is a metric, not a currency
At GTC 2026, Jensen Huang framed AI factories as power-limited systems. Every facility operates under a hard ceiling in gigawatts, and at a fixed power level, whoever achieves the highest throughput per watt has the lowest production cost. Tokens per watt is therefore the metric that governs an operator’s revenue capacity.
This is correct, and it is an operating metric. It is not a unit of account.
Every unit fails at one end or the other
Plot each candidate unit against two things a contract needs at once: whether it can be metered precisely, and whether the buyer can plan a budget from it. The units that are easiest to measure are the ones that mean least, and the unit that means most is the hardest to bound.
This is why both extremes lose the deal. Raw token billing is exact and unintelligible; pure outcome pricing is intelligible and unbounded. The settlement is a unit the supplier meters precisely and the customer budgets simply.
Pricing research increasingly recognizes the limits at both ends of the spectrum. As Bessemer’s pricing work documents, raw token pricing is technically measurable but unintelligible to a nontechnical buyer, who cannot forecast a budget from it. Pure outcome pricing aligns value but transfers uncontrolled cost and scope-definition risk to the supplier. Neither extreme survives an enterprise procurement cycle intact.
The practical settlement is a governed usage unit whose consumption varies by workload class and execution policy. The customer receives a stable contractual framework. The supplier retains responsibility for the underlying cost model, routing, and capacity optimization. That is a division of risk both sides can actually sign.
What tokenized deposits actually license
Institutional tokenization is real, and it is the most frequently misappropriated precedent in this market. Tokenized real-world assets are a measurable and growing category, tracked live by aggregators such as RWA.xyz, and major banks have moved settlement infrastructure into production, with J.P. Morgan’s Kinexys among them.
The BIS describes tokenized deposits as digital representations of commercial bank money on programmable platforms. They confer a direct claim on the issuing bank. They are redeemable at par for central bank money of the same currency, with finality. The BIS position is that tokenization should be built into the existing two-tier system anchored in central bank money, precisely because singleness and elasticity are what keep the system stable under stress.
Those properties are the instrument. Par redemption, bank balance-sheet status, supervision, and settlement liquidity are not decorations on a programmable ledger. They are the reason it functions as money.
They share design ideas. They share no substance.
The overlap is real and it is small. Everything that makes a deposit money sits outside it, and everything that makes a usage credit a service entitlement sits outside it on the other side.
Usage credits share selected design principles with permissioned tokenized systems. They are not structurally identical to tokenized deposits and should not be marketed as monetary instruments.
The broader lesson is the one worth taking. Programmability becomes more valuable when it is paired with strong identity, governance, legal rights, controlled access, and reliable settlement. That convergence is the relevant precedent. Financial tokenization for its own sake is not.
When issuance replaces revenue
Decentralized physical infrastructure networks offer a genuine warning, and it is worth stating precisely rather than dismissively.
Render Network’s published burn-mint documentation describes the mechanism plainly. Customers convert fiat to RENDER and the tokens are burned on completion of work. Node operators, however, are compensated through scheduled emissions allocated epoch by epoch: 9,126,804 RENDER in year one and 5,905,580 in year two, on a predefined schedule.
The structure is deliberate and openly documented. It also means supplier compensation and customer payment are two separate flows. When the emission schedule is the larger of the two, reported network activity may not represent sustainable commercial demand. A tradable token additionally creates a second source of demand, speculation, which is independent of compute utilization entirely.
None of this proves that every decentralized compute model is structurally invalid. Render, Akash, Bittensor and io.net differ in workload, supply mechanism, token design, customer base, and revenue disclosure. Emissions, lease spend, market capitalization, provider payouts, and protocol revenue are not interchangeable metrics, and comparisons that mix them are meaningless.
Two flows that never meet
Render’s documentation describes both sides plainly. Customers convert fiat and the tokens are burned. Node operators are paid from a separate, predefined emission schedule. Nothing links the size of one flow to the size of the other.
The structure is deliberate and openly documented. The point is not that it is hidden — it is that reported activity on such a network measures issuance as much as demand. Source: Render Network, Burn-Mint Equilibrium.
The useful output is a set of questions any compute-token network should be able to answer from public disclosure. A network that cannot answer them has told you something.
Who funds the suppliers?
What portion of supplier compensation comes from customer payments rather than issuance?
Same period, same basis?
Are revenue and emissions measured over matching periods and valuation bases?
What survives without incentives?
What utilization would remain if token rewards stopped tomorrow?
Can you buy without holding?
Can a customer purchase the service without taking on a volatile asset?
Is quality enforceable?
Are availability, quality and dispute obligations contractually binding?
The governed workload plane
Everything above is diagnosis. This is the constructive claim.
The opportunity is not to become another GPU marketplace or another crypto network. It is to become the governed workload plane through which institutional customers convert approved data, approved models, and purchased capacity into accountable work.
The atomic object is the Work Order. A Work Order connects the customer institution, the authorized actor, the research or business purpose, the source manifest, the workflow type, the execution policy, the budget authorization, the estimated reservation, the actual settlement, the output location, the status, and the evidence trail. The usage credit is simply the budgeting and settlement unit attached to that object.
The Work Order lifecycle
Five stages, and a division of responsibility at each one. What sits above the spine is contractible. What sits below it is the operating system, and it stays behind the boundary.
The boundary is the design. It makes the business model explainable without making the operating system reproducible.
What the customer sees are workloads, projects, grants, datasets, approvals, and outputs, plus execution policy classes, the estimated reservation against actual settled usage, and the evidence trail. What stays behind the boundary is provider selection, routing logic, cache strategy, batching, scheduling, model combinations, and workload-placement algorithms.
For a research university, this is the difference between a token wallet and an institutional system. A Work Order records the grant or project, the purpose, the actor, the data policy, the execution class, the source manifest, and the maximum credit budget. Required approvals (data-use restrictions, IRB, FERPA, HIPAA, CUI, export control, or sponsor conditions) are checked before execution rather than reported afterward. The institution buys accountable research capacity. It does not need to know which provider token, accelerator, or routing path produced each intermediate step.
What a governed unit must account for
A stable public contract requires a detailed private cost model. The two are not in tension; the first depends on the second.
The execution-cost layer has to include reserved-capacity cost, hardware depreciation or lease expense, power and cooling, networking, storage, runtime operations, security, model serving, orchestration, observability, support, expected utilization, and the treatment of failed work. Token processing is one component of that cost. It is nowhere near the whole of it, which is exactly why token counts make a poor commercial unit.
Four axes, and the unit they produce
Four of the five dimensions are independent inputs; the fifth is the unit they resolve into. Training versus inference and in-house versus approved-external vary separately, and collapsing any of them into another destroys the ability to price. The customer sees only the result.
The disclosure boundary is straightforward: a customer is entitled to the rate schedule and settlement rules required to understand their charges. A supplier is not obliged to disclose proprietary routing, model-combination, caching or capacity-allocation mechanics.
The evidence test
This market runs on claims that sound like facts. Here is a test that separates them, and it applies to every vendor in the category — including the one publishing this essay.
Six distinct things get described with the same confident verb tense: a market thesis, an approved architecture, implemented software, a deployed service, a contracted commercial position, and measured production performance. They are not equivalent, and the distance between the first and the last is where most of the disappointment in enterprise AI lives.
For any claim of the form “the platform does X,” ask what evidence would exist if it were true.
Six things, one verb tense
A thesis, an architecture, working software, a running service, a signed contract and a measured result all get described with the same confident “we do this.” They are not the same claim. The first four are drawn; the last two are built.
The convention is the same one used on this essay’s cover: what is only drafted appears as blueprint linework, and what actually exists is copper.
Claim, required evidence, and the language that is safe without it
Take this to a procurement call. Every row is answerable in a single meeting, and the third column is what a careful vendor says before the second column exists.
| The claim | Evidence that would exist if true | Safe language before proof |
|---|---|---|
| A no-bypass gateway is implemented | Active release, service path, auth and policy tests, negative bypass tests that fail closed | Designed as a no-bypass gateway |
| Credits settle through an audited ledger | Stored transactions, reconciliation, concurrency tests, defined failure treatment, audit export | Architecture defines reservation and settlement |
| The vendor owns or controls capacity | Executed ownership, lease, reservation or operating agreement | Planned capacity model |
| A blended rate is economically validated | Capacity model, utilization, token mix, fully loaded cost, sensitivity analysis | Illustrative planning rate |
| The platform protects margin from infrastructure volatility | Contract terms, supplier costs, routing limits, utilization and modeled downside cases | Designed to manage provider and workload mix |
| The platform is production-ready | Active release, security testing, recovery, observability, customer acceptance, stability evidence | Prototype, staging or production pilot |
The minimum proof package is short enough to request in a single call. An active deployed release and service path. Gateway policy and fail-closed negative tests. Hold, execute, settle, refund and concurrency tests. Tenant-isolation and authorization evidence. A reconciled usage ledger with an exportable audit record. Measured throughput, latency, utilization and fully loaded cost. And executed contracts, or planning assumptions clearly labeled as such.
Architecture documentation is not evidence of production readiness.
The thesis, and what would falsify it
Compute stays heterogeneous
AI compute will not converge on a single universal token. Physical capacity, model processing, enterprise usage rights and financial settlement remain distinct instruments with distinct issuers and obligations.
Governance gains value as supply fragments
A control plane is worth more when there is more to control between. Fragmentation is the condition that creates the layer, not the risk that destroys it.
The unit is the Work Order
The durable commercial object binds identity, policy, budget, data, execution, output and evidence. The credit is only the settlement unit attached to it.
Proof precedes the claim
Thesis, architecture, software, deployment, contract and measured performance are six different things. Only the last two are evidence.
What the thesis claims: that governance and settlement become more valuable as compute and models fragment; that capital intensity creates real risk for infrastructure owners even during strong demand; that institutional buyers need stable commercial abstractions and auditability; and that routing and workload control can create switching and policy value.
What the thesis does not claim: that all model tokens are fungible; that every NeoCloud fails or every hyperscaler wins; that usage credits are money, deposits, stablecoins or investments; or that any single broker’s throughput proves the ultimate margin structure of the category.
And because a thesis that cannot be wrong is not worth holding, here is what would falsify it.
- F1Supply consolidates. If a single provider achieves sufficient dominance that heterogeneity collapses, the control layer loses its reason to exist. Governance is worth less when there is nothing to govern between.
- F2Model quality converges. If routing becomes genuinely commoditized because every model performs equivalently, the switching value evaporates.
- F3Buyers accept raw token billing. If enterprises demonstrate durable willingness to consume without governance overlays, the commercial abstraction is unnecessary.
- F4Regulators standardize the unit. A mandated compute-accounting standard would preempt the private governed unit.
- F5Hyperscalers ship native governance. If platform-native controls satisfy institutional audit requirements, the independent layer is absorbed.
None of these look likely today. All of them are observable, and each would be visible well before it was decisive. That is the point of writing them down.
The most important shift underway is not from dollars to tokens. It is from ungoverned consumption to governed work. The winning institutional platform will not merely meter compute. It will connect identity, policy, budget, data, execution, output and evidence into one enforceable transaction. Everything else in this market is a supplier to that transaction.
Questions
Are model tokens and crypto tokens the same thing?
Why not just price AI work in GPU-hours?
Is tokens per watt a bad metric?
If inference prices are falling five to tenfold a year, why are AI bills rising?
Does CoreWeave’s operating loss mean GPU rental is a bad business?
What is a Work Order?
How is a usage credit different from a tokenized deposit?
What is wrong with paying suppliers in token emissions?
Why should routing and model selection stay proprietary?
How do I evaluate whether a governed-compute platform is real?
Read next
Sources
Standalone sources pageEvery source on one page — grouped, linked, citableSources are categorized by evidence type. Market prices and company metrics are point-in-time as of 11 August 2026 and should be refreshed before reuse. Preprints are research inputs, not peer-reviewed authority unless a publication venue is separately confirmed.
- CoreWeave — Q1 2026 resultsPrimary company disclosure. investors.coreweave.com
- Gundlach, Lynch, Mertens & Thompson — The Price of ProgressResearch preprint, arXiv:2511.23455 (MIT FutureTech). arxiv.org
- Epoch AIInference price and capability datasets. epoch.ai
- Artificial AnalysisModel pricing and performance benchmarks. artificialanalysis.ai
- Menlo Ventures — OpenRouter token run rateInvestor source. menlovc.com
- Mobile World Live — NVIDIA on AI factories and tokensTrade press. mobileworldlive.com
- Bessemer Venture Partners — AI Pricing and Monetization PlaybookIndustry analysis. bvp.com
- Bank for International Settlements — Anchoring Trust in MoneyAnnual Economic Report 2026, Chapter III. Primary institutional source. bis.org
- Render Network — Burn-Mint EquilibriumPrimary protocol documentation. know.rendernetwork.com
- SemiAnalysis — GPU Cloud IndexMarket index and industry research. semianalysis.com
- Andreessen Horowitz — Welcome to LLMflationIndustry analysis. a16z.com
- Deloitte — Compute Power and AI2026 TMT Predictions. Industry forecast. deloitte.com
- J.P. Morgan — KinexysPrimary company source on tokenized settlement. jpmorgan.com
- RWA.xyzLive tokenized real-world asset market data. rwa.xyz
- NPR Planet Money — AI tokens and the utility analogyJournalism. npr.org