You type a sentence, wait 50 milliseconds, and an answer streams back. To most users the whole thing feels weightless. Then the invoice arrives: millions of tokens in, millions out, context length, cache hit or miss. One thought flashes: an intangible reply, metered like a utility — why?
The short answer is that you are looking at three different bills at once, and mixing them together is what makes AI pricing feel like a contradiction.
Three bills, one answer
- Production cost — what it actually takes to mint one token: electricity, chips, data centers, engineering.
- Market price — what the vendor quotes per token, which already includes competition, profit and tiering.
- Task cost — how many tokens a whole job burns, end to end.
Confuse them and you get the two most common complaints: "electricity is cheap, why is AI so expensive?" and "prices keep dropping, why did my bill grow?"
Software is copyable at near-zero marginal cost. Inference is not. Every generation runs through real silicon, real memory, real power and real racks. "A token looks like software," the book Token Economy argues, "but it is increasingly heavy industry."
The five-layer cake
In January 2026 at Davos, NVIDIA CEO Jensen Huang described AI infrastructure as a five-layer cake, bottom to top: Energy, Chips, Infrastructure, Model, Application — a frame he repeated at CES and GTC 2026. Here is where the money actually lands, for large-scale commercial inference:
- Energy (10–20%). Real but not dominant. At US industrial rates ($0.06–0.08/kWh), a 500MW data center burns over $300M a year in electricity. But once GPU, HBM, network and facility depreciation are counted, energy is only 10–20% of a token's fully loaded cost — the widely shared "electricity myth" confuses operating cost with total cost of ownership.
- Chips (40–55%). The biggest layer and the supply bottleneck. A B200 sells for $30k–40k, and HBM plus advanced packaging is roughly two-thirds of a high-end chip's bill of materials. The technical reason memory decides supply: model weights must live in HBM for efficient inference, so HBM capacity and bandwidth cap speed and concurrency. HBM is an oligopoly — SK hynix ~57%, Samsung ~22%, Micron ~21% — and memory's share of hyperscaler capex jumped from ~8% (2023–24) to ~30% (2026), almost all HBM-driven.
- Infrastructure (15–20%). Turning thousands of chips into a working "AI factory": power delivery, liquid cooling, racks, optics, InfiniBand fabrics, orchestration. The share swings wildly with GPU utilization — the same silicon at 60% vs 80% utilization produces tokens at very different unit costs.
- Model (10–15%). The softest layer and the fastest to fall. Training is enormous — GPT-4 cost ~$40M in compute-only terms, over $100M per Altman's public figure, and GPT-5-class runs into the billions to tens of billions, before the trial-and-error factor. But architecture tricks cut cost without new hardware: DeepSeek V4 has 1.6T parameters yet activates only ~49B (≈3%) per token — a 16,000-person consultancy sending a ~500-person strike team to each engagement. Distillation lets a small model reproduce a large one's capability at ~1% of the inference cost.
- Application (5–10%). Interface, integration, compliance and content safety — what turns capability into something a customer will actually buy.
The exact percentages shift with your setup — owned vs rented, frontier vs economy models, training vs inference, long vs short tasks — but one ordering is stable: the three hardware-related layers dominate, typically around 70% or more.
Two myths, one "double clock"
Myth 1: a token is just resold electricity. The marginal electricity of one reply is a fraction of a cent. But a token's fully loaded cost also carries chip depreciation, memory-bandwidth opportunity cost, data-center utilization and a service-quality premium. Same kilowatt-hour, different chip, different utilization — wildly different number of tokens.
Myth 2: per-token cost is the metric. The same nominal token is not the same cost. Frontier flagship inference can be tens of times pricier than economy tiers; long-context tasks eat far more memory than short Q&A; and batching lifts cluster utilization from 60% to 70–80%, amortizing cost — which is exactly why Google's 2026 tiered inference pricing splits the same model into priority, standard, flexible and batch, with batch at half price — and DeepSeek's weekend off-peak pricing is the same idea: time-of-use pricing is becoming the industry default.
Underneath both myths sits the book's central idea, the double clock: software moves fast, physics moves slow.
- Fast clock (algorithms). Equal-capability inference cost fell ~99.5% in three years (to about 1/250, per Stanford's 2026 AI Index). MoE, attention optimization, distillation and speculative decoding iterate in days and weeks, on hardware you already own.
- Slow clock (physics). Data centers take 18–36 months and are slipping; of ~16GW of planned US capacity for 2026, only ~5GW was under construction by April. Transformers average a 128-week lead time (vs 6–8 months pre-pandemic) yet are under 10% of build cost — a $2B campus can sit idle for a year over a $40M transformer order. HBM capacity grows 30–50% a year against doubling demand. And electricity: IEA projects global data-center power to double from ~415TWh to over 900TWh between 2024 and 2029; Goldman sees a 45GW US deficit by 2028; the US interconnection queue sits at 2060GW with a median wait near five years — Google reports some sites queued for 12.
OpenAI's Stargate is the emblem: $500B announced with great fanfare, construction paused on grid queues and transformer delays. As the Financial Times put it, "$500 billion cannot buy speed from the physical world, nor consent from the community" — the same dynamic behind America's first lost summer for data centers, where 550 restriction bills made consent the scarcest compute resource.
Think of it like steel and cars: steel prices explain the steel, not why the car costs what it does. The algorithm clock explains why prices crash; the physical clock explains why scarcity never goes away.
Why this matters
The double clock resolves the two facts that seem to contradict each other: AI price wars fight to the bone while compute stays tight, and hyperscalers everywhere are rebuilding their balance sheets like infrastructure companies. A meter of this makes sense only if you know that every light-looking answer is minted from real energy, real chips and real rack space — and that those layers move on a clock measured in years, not weeks.
When intelligence becomes a metered, priced, callable commodity, the competition stops being "who writes better code" and becomes "who can organize energy, chips, models and human demand most efficiently." For model companies, that means converting every watt and every GPU into more high-value tokens. For clouds, it means keeping silicon and racks busier. For applications, it means proving tokens actually produced an outcome someone pays for.
What to do about it
- Separate the three bills. When you negotiate or budget, know whether you're looking at production cost, list price, or whole-task cost — mixing them produces bad decisions.
- Buy the slow lanes. Adopt batch and flexible tiers (half price at Google's 2026 batch tier) for workloads that can wait; reserve priority lanes for what must be instant.
- Watch utilization, not just price. The same chips at 80% utilization can halve your effective per-token cost versus 60% — schedule bursty work into the gaps.
- Plan for the 12–18 month lag. NVIDIA's hardware cuts per-token cost 40–60% per generation, but a chip announced today ships at scale a year or more later. Cheap compute at a specific place and time is still scarce.
- Track HBM, not just GPUs. Memory is the new binding constraint — memory price hikes just pushed AI servers up over 15%, shifting pricing power from GPU vendors to memory makers — and its share of cloud capex tripled in two years. Allocation, not raw FLOPs, is becoming the real queue.
The takeaway is not a single number — there is no "dead price" for a token. Prices change, models change, cost structure changes with task and deployment. But one thing is fixed: every token is produced by real energy, chips and infrastructure. That is the real starting point for understanding the AI economy.
FAQ
Why do AI prices keep falling while my bill keeps growing?
You are mixing the three bills: vendors cut the market price, but your invoice tracks task cost — growing usage, longer contexts and frontier-tier upgrades raise total token burn, while physical layers (chips, data centers) fall on a 2-5 year clock and scarcity persists.
Which layer of the five-layer cake deserves the most attention?
Chips (40-55%) are the cost backbone, and within them HBM is the supply bottleneck: SK hynix ~57%, Samsung ~22% and Micron ~21%. Capacity grows 30-50% a year against doubling demand, and memory's share of cloud capex jumped from ~8% (2023-24) to ~30% (2026).
What can teams do to cut AI costs today?
Route wait-tolerant workloads to batch or flexible tiers (half price at Google's 2026 batch tier); push utilization higher — 80% vs 60% can nearly halve effective unit cost; and budget for the 12-18 month hardware lag, since cheap compute at a given place and time is still scarce.