All research sources
Sources & Citations

The Inference Paradox

23 primary sources 8 categories We publish the receipts
The essay The Inference Paradox
A

Cheaper tokens, costlier work

Gartner · The Register

The cost of running an agentic AI workflow will rise more than fivefold through 2028.

Gartner, Aug 17, 2026 View source

In the same breath, Gartner expects the per-token price of frontier inference to fall by more than 90 percent by 2030.

Gartner, Mar 25, 2026 View source

The cost of the thing you actually buy, a completed workflow, goes up.Will Sommer, the Gartner analyst behind the forecast, said it plainly: routing a task to an agentic reasoning model increases inference costs at least fivefold, and…

The Register View source
B

Four terms up, one term down

Computerworld

It is what you get when four things compound upward and one thing compounds downward, and nobody in the building owns the ratio. 02 · The paradox in one equation ## Four terms up, one term down Gartner built its forecast on a Tokenomics…

Computerworld View source
C

Reasoning tiers cost an order of magnitude more. Sometimes two.

NVIDIA Q4 FY2025 · DeepLearning.AI · ValueAddVC

Sometimes two. Jensen Huang said on NVIDIA's fourth-quarter fiscal 2025 call that reasoning can demand 100 times more compute than a simple query, that the vast majority of NVIDIA's compute today is actually inference, and that future…

NVIDIA Q4 FY2025 call, transcript View source

Benchmarking GPT-4o cost $109. o1-pro and GPT-4.5 were priced at $600 per million output tokens against $15 for Claude 3.5 Sonnet.

DeepLearning.AI, The Batch View source

On a medium-complexity coding prompt, o1 ran at $0.450 per request and o3 at $0.364 against $0.011 for GPT-4, roughly 33 to 40 times more.

ValueAddVC View source

Even a cheap reasoning model loses its headline advantage once the reasoning tokens are counted: a single DeepSeek R1 request can generate 5 to 50 times the tokens of V3.1 (DeployBase). 02 ### Agents burn more tokens than chat.

DeployBase View source
D

Agents burn more tokens than chat. Far more than Gartner's range.

Anthropic · Microsoft Research · Spheron

Anthropic also found that token usage by itself explains 80 percent of the variance in performance on the BrowseComp evaluation.

Anthropic Engineering View source

Runs on the same task differed by up to 30 times in total tokens, and higher token usage did not translate into higher accuracy (Microsoft Research).At the workflow level the same thing shows up in ordinary examples.

Microsoft Research View source

Seventeen times more for what the user experiences as one question answered.

At 30 turns the actual cost is 15.5 times the naive linear estimate (n1n.ai). 03 ### Deflation is real, fast, and resets at every new tier. The deflation half of the paradox is well documented and in places runs faster than Gartner says.…

E

Deflation is real, fast, and resets at every new tier.

a16z · Epoch AI

The same piece notes that OpenAI's o1 launched at the same $60 per million output tokens that GPT-3 charged at launch (a16z, LLMflation).Epoch AI tracked the price of reaching a fixed capability across six benchmarks and found declines…

a16z, LLMflation View source

The same piece notes that OpenAI's o1 launched at the same $60 per million output tokens that GPT-3 charged at launch (a16z, LLMflation).Epoch AI tracked the price of reaching a fixed capability across six benchmarks and found declines…

Epoch AI View source

Epoch's read of the trend is very roughly a 5 to 10 times cost reduction per year for reaching a given capability level (Epoch AI, Feb 2026).So the price of yesterday's capability collapses.

Epoch AI, Feb 2026 View source
F

The tier-mix shift is the quiet driver, and the scaling law pushes it.

OpenAI · Snell et al.

OpenAI wrote it down in the o1 announcement: the performance of o1 consistently improves with more reinforcement learning and with more time spent thinking.

Snell and colleagues at UC Berkeley and Google DeepMind formalized it as a scaling law: scaling test-time compute can be more effective than scaling model parameters (Snell et al., 2024).When the law says more inference compute buys…

Snell et al., 2024 View source
G

At the industry level, demand is outrunning efficiency.

SemiAnalysis · SemiAnalysis via · IDC

SemiAnalysis models AI demand at 86 percent of TSMC's 2027 N3 wafer output, nearly entirely squeezing out smartphone and CPU wafers, with 2026 hyperscaler capex consensus moving materially higher and Google's roughly doubling.

SemiAnalysis, Mar 2026 View source

Meta raised 2026 capex guidance to $125 to $145 billion, nearly double its $72.2 billion in 2025; combined 2026 capex for Meta, Microsoft, Alphabet and Amazon is estimated around $725 billion, up 77 percent; Meta signed more than 5 GW of…

SemiAnalysis via Bitget, Jul 2026 View source

IDC expects AI spending to grow 31.9 percent a year to $1.3 trillion in 2029, more than 26 percent of worldwide IT spending (IDC, Aug 2025).The chips are getting cheaper per operation.

IDC, Aug 2025 View source
H

The failure signature is already in the survey data.

Gartner · MIT NANDA · S&P Global Market

Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, naming escalating costs, unclear business value and inadequate risk controls.

Gartner, Jun 25, 2025 View source

MIT's NANDA group found that only about 5 percent of enterprise generative AI pilots achieve rapid revenue acceleration, based on 150 interviews, a 350-person survey and 300 public deployments; purchased tools succeeded about 67 percent of…

MIT NANDA, via Fortune View source

S&P Global surveyed more than 1,000 enterprises and found the share abandoning most of their AI initiatives before production jumped from 17 percent to 42 percent in a year, with 46 percent of proofs of concept scrapped before broad…

S&P Global Market Intelligence View source
These are the primary sources behind our research. Have a better one, or spot an error? Tell us — we’ll correct it.