Token Economics: The Unit of Account Rewriting AI

Most people first think seriously about the cost of AI when an invoice arrives: how many tokens did we burn this month? Almost nobody stops to ask the deeper question — what does a token actually measure?

The answer is hiding in economic history. Economists have a blunt method for identifying an industrial revolution: look at its unit of account. In the agricultural age, the bushel made it possible to buy Kansas wheat in Chicago sight unseen — and futures markets followed. The oil barrel, standardized at 42 gallons in 1866, turned crude from a local mineral into an asset that could be hedged, shorted, and traded globally. The kilowatt-hour, first proposed by British engineers in 1889, gave electricity a price scale. The bit, defined by Claude Shannon in 1948, turned bandwidth and storage into commodities with a price tag.

These four units share one property: they all measure something reducible to a physical process — matter, volume, energy, the encoded length of a signal. What a bit stores, or whether it is useful, is outside the concept entirely.

The token breaks the pattern. It does not measure matter or energy. It measures the output of a machine performing an intellectual task. Understanding a sentence, generating code, making a judgment — activities that once existed only in human brains — now have a countable unit and an executable price for the first time in history.

That is what a token really is: not a billing detail, but the bushel of the intelligence-as-a-service era. And this new unit of account is quietly pulling out the foundation under some of economics' most stable assumptions.

1. The First Unit of Account You Cannot Price Upfront

A token has a property none of its predecessors had: its value is determined entirely after the fact, by the task it was spent on.

The quality of a bushel of wheat was fixed at harvest. The grade of a barrel of crude was measurable at the wellhead. A token is the opposite. Using the same model to review a million-dollar M&A contract and to chat about the weather can consume roughly the same number of tokens — while producing economic value orders of magnitude apart. And that difference only becomes visible after the call completes. It cannot be labeled or graded in advance.

Worse, the marginal cost of switching between uses is close to zero. Same model, same API, one changed instruction: from small talk to contract review. The production side is uniformly measurable; the consumption side delivers highly non-standard intellectual output. The two are completely decoupled.

Supply-and-demand analysis rests on a default assumption: production characteristics and consumption value are coupled. That assumption no longer holds. You pay for "ten million tokens," but what you bought might be lawyer-grade work product — or half an hour of weekly-report drafting.

This is why the standard toolkit starts failing. Your legal team burns 50 million tokens reviewing contracts this month and triple that next month. Your engineering team deploys a few agents and consumption jumps 10x in a week, with no procurement model able to predict which department, when, or how much. The CFO asks for an "enterprise token ROI" and discovers no consistent denominator exists.

2. Unit Prices Collapse, Total Spending Explodes: The Jevons Clock

Intuition says a commodity whose price is in freefall cannot become a big deal. The data says the opposite.

At GPT-4-level capability, inference cost roughly $60 per million output tokens in March 2023. By 2025, equivalent capability sold for around $1 — third-party trackers such as Epoch AI's inference price dataset show median prices for matched benchmarks falling by roughly two orders of magnitude per year. Yet over the same period, total AI infrastructure spending went up, not down: IDC put full-year 2025 AI infrastructure spend at about $318 billion, more than double 2024's roughly $153 billion.

Demand is even more dramatic. At DevDay in October 2025, OpenAI reported its API processing about 6 billion tokens per minute. By March 2026, the figure exceeded 15 billion per minute — a 1.5x jump in five months, excluding all ChatGPT consumer traffic. China's trajectory is steeper still: the National Data Administration disclosed that daily token consumption climbed from roughly 100 billion in early 2024 to 100 trillion by the end of 2025, and past 140 trillion by March 2026 — a more-than-1,000-fold increase in two years. Inference now accounts for roughly two-thirds of all AI compute, up from about a third in 2023.

This is the Jevons paradox, version 2.0. In the 19th century, more efficient steam engines consumed more coal, not less. Today, model prices fall more than 90% a year, and total token consumption inflates even faster. The structural reason is clear: algorithmic efficiency improves independently of hardware, but a fab takes three to four years from groundbreaking to production. Two clocks running at completely different speeds — and the price-collapse clock will always outrun the supply-expansion clock.

There is a subtler variable too: the decision to consume tokens is shifting from humans typing prompts to agents triggering calls continuously in pursuit of goals. When agents decide which model to call and how many tokens to spend, demand stops responding to price and latency the way any familiar consumer market does. The agent-interoperability push by device makers and super-apps is, in essence, plumbing for this automated consumption machine.

3. The Three Layers of the Token Market: A Framework

Assemble the pieces above and a reusable three-layer structure emerges.

Layer 1: The Metering Layer (uniform). All intellectual output is compressed into one unit — the token. This layer is highly standardized; competition happens on price and latency. This is the model vendors' battlefield: price per million tokens, throughput, context window.

Layer 2: The Value Layer (heterogeneous). The same tokens entering different tasks produce economic value orders of magnitude apart, invisible upfront. There is no market-clearing price here, only ex-post ROI reviews. This is the application companies' battlefield: whoever steers tokens toward high-value tasks creates 10x the value at the same cost.

Layer 3: The Trigger Layer (automated). Who decides to consume? The answer is shifting from humans to agents. This layer shapes demand itself — fragmented, volatile, and compounding. This is the platforms' battlefield: operating systems and super-apps are fighting over the invocation entry point for agents.

The framework explains three puzzles. Why price wars never kill model vendors: metering-layer margins are thin, but layers two and three settle their profits elsewhere. Why enterprise AI procurement feels broken: procurement processes were designed for the metering layer, while the value is generated in the second. And why every platform is racing to build agents: the trigger layer is the demand amplifier — whoever holds trigger rights holds total token volume.

4. Applications: From Jensen Huang's Ledger to Your Department Budget

At GTC in March 2026, Nvidia CEO Jensen Huang offered a deliberate number: an engineer earning $500,000 a year who consumes less than $250,000 in tokens would worry him deeply. If a top engineer said "I plan to work with pen and paper only," he added, that would be as absurd as a chip designer refusing EDA tools. Nvidia has even started offering token allowances worth nearly half an engineer's salary as a hiring incentive.

The right way to read this is not to memorize $250,000 — a number that will date quickly — but to see the emerging standard behind it: token consumption is becoming a new dimension for measuring knowledge-worker productivity, the way bandwidth and cloud specs once measured a company's digital maturity.

Run the three layers forward:

For model vendors, the metering layer is a red ocean. Prices collapsing toward marginal cost is irreversible; differentiation must come from the layers above — which is why OpenAI and Anthropic push enterprise offerings and agent products. They are hunting profits in layers two and three.

For application companies, the value layer is everything. Jev, the classification model TypeSafe released in September 2026, is an extreme specimen: it writes nothing, explains nothing, chats about nothing — it just picks one option from a list, at under $0.001 per call. Narrowing a model to a single high-value task is how you push the value layer's signal-to-noise ratio to its limit while metering-layer costs stay fixed.

For platforms, the trigger layer is the endgame. The hardware response tells the same story: the "AI box" category — with Apple, Nvidia, and AMD all fielding entries — packs over 100 GB of unified memory into a desk-side enclosure. The point is to move token consumption from the cloud's metering layer back on-premises: data never leaves the room, but every token is still metered.

5. What To Do About It

If you manage a business:

  • Stop asking "which model should we buy." Ask first: "where is our list of high-value tasks?" Metering-layer vendors will keep changing; the value-layer task definitions are the asset.
  • Manage token budgets like R&D budgets, not IT procurement. Demand is inherently fragmented and volatile; annual plans will miss. Quarterly rolling budgets with per-team allocations are more realistic.
  • Track two numbers: token-consumption growth per team, and per-task value reviews. The first shows where agents are penetrating; the second shows which segment of the value layer your money actually reached.

If you are a developer or individual practitioner:

  • Use token consumption to calibrate your career: if your tasks cannot absorb large volumes of high-value tokens — highly repetitive, low judgment density — that role gets metered first.
  • Learn to split work into "metering-layer outsourceable, value-layer self-owned": hand generation, translation, and first drafts to tokens; keep judgment, taste, and accountability for yourself.

Every rewrite of economic order comes with a new unit of account. This time, for the first time, what is being measured is intelligence itself.

Seeing tokens not as payment for compute, but as the price of "one machine-performed act of intellectual work," is the entry ticket to the intelligence-as-a-service economy. The winners will be those who build their own value layer on top of the unit everyone else merely pays for.

Disclaimer: This article is for informational purposes only and does not constitute investment advice.

Scroll to top