Token Economics: The First Meter for Intelligence

Every AI pricing page reduces to the same unit: dollars per million tokens. Most people read that line as a billing rule — a cost item to be optimized away. That reading is too narrow, and it quietly misprices the biggest economic shift of this decade.

The argument of this piece is simple: the token is the first unit in history that lets us buy, meter, and bill the output of intelligence itself. That is not a pricing detail. It is a structural event, because nearly every analytical habit we bring to tokens was built for oil, electricity, or bandwidth — and those frameworks fail against tokens in predictable, compounding ways.

Every Economic Era Gets the Unit It Deserves

Economists have a crude but reliable test for whether an industrial revolution is real: did it produce a new unit of account?

Agriculture had the bushel. Before standardization, grain was priced by the sack, and a sack's "quantity" depended on the sack and the seller's honesty. The bushel made it possible for a buyer in Chicago to purchase Kansas wheat he had never seen — remote futures trading existed because the unit did. The railroad cut transport costs, but the bushel turned grain into a globally tradable commodity.

The energy era had the barrel. The 42-gallon standard, fixed in 1866, turned oil from a Pennsylvania curiosity into an asset class that could be hedged, shorted, and futures-traded in New York and London.

Electrification had the kilowatt-hour (kWh). Proposed by British engineers in 1889 and formally recognized by the International Electrotechnical Commission (IEC) in 1948, it is what made grid dispatch, power markets, and industrial energy management possible at all.

The information age had the bit. Claude Shannon's 1948 paper, "A Mathematical Theory of Communication," turned information from a philosophical abstraction into a measurable quantity — and line capacity, disk space, and bandwidth all acquired price tags.

These four units share one property: what they measure reduces to a physical process. The bushel measures matter, the barrel volume, the kWh energy, the bit the encoded length of a signal. Whether the content of a bit is useful is outside the bit's scope entirely.

The token breaks that pattern. It measures none of these things. It measures the output of a machine performing an intellectual task — understanding a sentence, generating code, reaching a judgment. Activities that used to exist only inside human heads, bundled into salaried hours and headcount, now have a meter and a price.

That is the material basis of intelligence-as-a-service. Buying intelligence used to mean buying a person, their hours, their role. Now it means buying tokens: task-shaped, result-shaped, callable on demand. For the first time, cognition has been detached from an individual brain and made continuously meterable.

Two Anomalies That Break the Old Frameworks

If tokens were just another unit, this would be a curiosity. What makes them structurally disruptive is that they violate two assumptions conventional commodity analysis depends on.

Anomaly one: homogeneous metering, heterogeneous value.

A bushel of wheat has a fixed quality at harvest; a barrel of crude has a measurable grade at the wellhead. A token is the opposite. Its value is determined entirely by the use case, and it cannot be graded in advance. The same model reviewing a nine-figure merger contract and the same model chatting about the weather may burn a comparable number of tokens while creating economic value orders of magnitude apart — and you only learn which after the fact.

Worse, the marginal cost of switching between those use cases is roughly zero. One API, one model, one prompt away from weather small talk to contract review. The result is a market structure with no real precedent: production is uniformly metered, but consumption delivers wildly non-standard intellectual output. Traditional supply-and-demand analysis assumes producer characteristics and consumer value are coupled. For tokens, that assumption is simply false.

Anomaly two: two clocks running at different speeds.

Using GPT-4-class capability as the benchmark, inference cost fell from roughly $60 per million output tokens in early 2023 to under $1 for equivalent capability by 2025 — nearly two orders of magnitude in two years. Yet global AI infrastructure spending rose over the same period: IDC puts 2025 spending at $318 billion, more than double the $153 billion of 2024.

Falling unit prices with exploding total spend is the Jevons paradox with a fresh coat of paint: when 19th-century steam engines became more efficient, coal consumption rose rather than fell, because efficiency unlocked new uses that consumed the savings. This time the loop spins faster, because the two halves of the market run on different clocks. Algorithmic efficiency improves month to month on existing hardware; a fab takes three to four years from groundbreaking to production. Software moves in months, silicon in years — and both clocks govern the same market.

The Economy Is Already Running

This is not a forecast; the token economy is operating now, and its slope is the story. Three data points:

OpenAI disclosed in March 2026 that its API was processing 15 billion tokens per minute — API traffic only, excluding consumer ChatGPT usage. In October 2025, at its developer day, the figure was 6 billion. A 1.5x increase in five months.

China's curve is steeper. The National Data Administration reported in March 2026 that daily token consumption had grown from roughly 100 billion in early 2024 to 100 trillion by the end of 2025, and past 140 trillion by March 2026 — a thousandfold expansion in two years.

The compute mix has flipped accordingly: inference now accounts for roughly two-thirds of AI compute, up from about one-third in 2023. Train once, infer billions of times is becoming the dominant shape of the spending curve.

At a GTC interview in March 2026, Jensen Huang framed what this means for individuals: "An engineer making $500,000 a year — if they consume less than $250,000 worth of tokens a year, I'd be very worried." He added that a top engineer insisting on working with pen and paper is as absurd as a chip designer refusing to use electronic design automation (EDA) tools. The dollar figures will date quickly as prices fall, but the standard they point to will not: token consumption is becoming a measure of how much machine leverage a knowledge worker is actually applying, the way bandwidth and cloud specs once measured a company's digital maturity.

Why Your Existing Playbook Fails

Put the structure into a concrete scenario and the failures appear immediately. Suppose you are a strategy lead asked to build your company's AI procurement function from scratch. Nearly everything an MBA curriculum teaches about supplier evaluation breaks:

On the supply side, dozens of model providers operate with fundamentally different cost structures — hundreds of billions in cumulative investment at the frontier, while open-source teams fine-tune near-frontier capability for a few million dollars. On the demand side, your legal team burns 50 million tokens reviewing contracts this month and might triple that next month; after deploying a few agents, engineering consumption jumps tenfold in a week, and nobody can predict which department, when, or how much. The stable annual demand plan that procurement is built on does not exist for tokens.

Pricing analysis fares no better: the assumptions behind competitive pricing models — homogeneous products, free entry and exit, stable marginal costs — all fail. You cannot even guarantee what capability next quarter's budget will buy. And value assessment is the hardest problem of all: the same 10 million tokens might let lawyers avert a seven-figure legal risk, or save marketing half an hour on a status report. Yet the CFO still wants a single, unified "enterprise token ROI."

One more variable is arriving: call decisions are shifting from humans typing prompts to agents triggering consumption continuously under standing goals. When agents decide which model to call and how many tokens to spend, demand stops responding to price and latency the way consumer markets do. Cheaper tokens pull agents into more workflows, agents amplify demand, demand exposes physical supply limits, and those limits rewrite prices and adoption thresholds. The loop is already turning.

What To Do About It

A framework earns its keep only in decisions. By role:

  • If you run finance or procurement: abandon the search for one token ROI. Bucket spend by task leverage — high-leverage work (contract review, code generation, customer analysis) versus low-leverage work (drafts, summaries) — and treat the former as investment, only the latter as cost. Build demand volatility into the budget model itself; the annual plan is dead.
  • If you manage engineering or product teams: treat token consumption as a productivity dashboard, not a cost ceiling. A sudden drop usually means AI usage is receding, not that you saved money; a spike may mean a workflow finally works. Read the tasks behind the tokens before judging the number.
  • If you build products: two orders of magnitude of price decline in two years means applications that fail unit economics today may clear them next year. Assume prices keep falling, and put your moat in data, workflow depth, and domain understanding — not in today's inference-cost arbitrage.
  • If you are a knowledge worker: read Huang's warning in reverse. Token consumption will increasingly measure your leverage — not how much you type, but how much machine intelligence you mobilize. Learning to translate tasks into high-quality context and instructions is the new baseline skill.

The bushel made wheat remotely tradable, the barrel made oil a financial asset, the kilowatt-hour made the grid operable, the bit gave information a price. The token is doing the same thing to intelligence itself. Understanding this unit is not about reading a pricing page. It is about reading the rewrite of an economic order that is already underway.

Scroll to top