Cloud Giants Take 40% of Every $100 AI Labs Earn

Every time an AI lab books $100 of revenue, between $35 and $40 of it flows straight to AWS, Azure, or Google Cloud as inference compute fees. That is the central finding of a Barclays unit-economics model released August 28 — and it explains the most important business reality of the AI boom: the landlord collects a cut of every tenant's rent before the tenant keeps a dime. It is the income-statement version of the compute-concentration trend we have been tracking.

Here is the full breakdown of where the money actually goes, why the labs' own margins have suddenly exploded, and why the clouds' grip starts slipping in 2028.

Inference margins just tripled — because agents became must-buy

The headline number inside the report is not the cloud cut; it is the speed of the margin recovery. Barclays estimates that paid inference margins at frontier labs have climbed from low-double-digit in 2025 to 50–65% or higher in 2026, with adjusted gross margins up 30–50 percentage points year over year.

The driver is demand structure, not pricing magic. Enterprise customers and agentic workflows have become "must-buy" products — the kind of workload that does not get cut when budgets tighten. That changes the pricing power of the labs more than any single price hike could.

But not every product line makes money equally. Barclays breaks the margin ladder into three tiers:

  • Direct API — 80%+ inference margin. The oldest and richest line. Developers in tools like Cursor and Figma pay per token consumed, and token efficiency gains, nominal price increases, and infrastructure improvements (quantization, speculative decoding, newer compute platforms) keep pushing the margin higher. Barclays notes Q2 2026 API margins were actually well above what its own published charts show.
  • Indirect API — same experience, different billing. The user experience is identical to direct API, but the billing relationship runs between the customer and the cloud provider. As this share grows, accounting gaps between labs widen and financial statements become harder to compare.
  • Subscriptions — ~70%, the lowest tier. Products like Anthropic's Claude Code and OpenAI's Codex deliberately subsidize token costs to defend retention. Fixed monthly fees with usage caps, plus increasingly frequent usage-limit resets, keep margins thinnest here — a visible tension between growth retention and unit economics.

The 17-point gap no investor should miss

To make the accounting mess concrete, Barclays built two hypothetical frontier labs. Lab A gets ~70% of revenue from API and ~30% from subscriptions; Lab B is the mirror image at 20% API / 80% subscription. After cost allocation and partner revenue-share adjustments, their adjusted gross margins differ by a whopping 17 percentage points — roughly 55% for Lab A versus 38% for Lab B.

The gap is not all real economics; some of it is bookkeeping. Lab A recognizes indirect API revenue on a gross basis; Lab B uses net accounting or skips partner-operated indirect revenue entirely. Barclays compares it to the classic Uber-versus-Lyft accounting gap: nearly identical core businesses, wildly different reported revenue and profit, purely because of how revenue is recognized.

The practical lesson: do not compare AI labs on reported revenue or margin one-to-one. A subscription-heavy lab will look structurally poorer on paper even when the underlying business is healthy.

The landlord's cut, in hard numbers

Now the cloud side of the ledger. For Lab A, every $100 of lab revenue puts $35 in cloud provider revenue, roughly $11.8 of profit, a ~34% operating margin. For Lab B, the cloud takes $41, with $19.1 of profit and a 47% operating margin.

But part of that cloud margin is an illusion created by the partner revenue-share agreements — roughly 20% of revenue, capped and cumulative, which Barclays expects to be phased out after 2028. Strip out the revenue share and the cloud's per-token profit is no different between the two labs.

There is also a second, quieter revenue stream: agent subscriptions pull in extra cloud spend. Agentic workloads persistently save task state and call databases, storage, and other cloud services — so every agent subscription sold adds cloud revenue beyond the inference fee itself. It is the same "sell apps for tokens, sell infrastructure for profit" layering we broke down in our token-economy analysis.

The window closes in 2028

AI lab revenue is scaling absurdly: from roughly $7B in 2024 to $137B in 2026 and an estimated $690B by 2028 (annualized recurring revenue reaching ~$200B by end-2026 and ~$782B by end-2028).

More importantly, the ratio that defines the clouds' power is collapsing. Cloud AI revenue stood at 153% of AI lab revenue in 2024 — labs were spending more on compute than they earned. That ratio falls to 90% in 2026 and 73% in 2028. Training spend as a share of lab revenue drops from 96% in 2024 to 48% today, and is projected at 35% by 2027 and 30% by 2028.

The mechanism behind the slide is physical: from 2028, dedicated compute infrastructure that labs have already signed and locked in starts coming online. AWS, Azure, and GCP hold roughly stable share of lab compute spend for the next two years — then the clouds begin losing ground in both training and inference to the labs' own capacity.

Barclays also warns the 50–65% inference margins will not stay this fat. Frontier-model competition is intensifying and compute supply keeps growing, which argues for margin compression over time. The landlord's 40% tax is real today — but it is a window, not a permanent constitution.

What to do with this

  • If you are a founder or platform engineer: inference unit economics are the real price of doing business. Evaluate vendors on cost per token, quantization, and speculative-decoding gains — not headline model price. The 40% cloud cut is a fixed tax until your own capacity comes online.
  • If you are an investor: adjust gross margin is the metric that matters, and it is only comparable after normalizing gross-versus-net revenue recognition. Treat any one-lab-versus-another margin comparison as suspect until you know which accounting basis each uses — the same trap we flagged in our two-economies analysis of cloud profits vs. application losses.
  • If you are an enterprise buyer: an agent subscription bill is not just token fees. Budget for the state, database, and storage spend that agentic workloads silently generate on the same cloud.
  • If you price AI products: direct-API margins above 80% tell you inference is a great business right now — but competition guarantees erosion. Design pricing for the 2028 world, where labs run their own compute and the cloud tax shrinks.

Source: Barclays AI industry unit-economics research (August 28), via Cailian Press / 36Kr.

Leave a Comment

Scroll to top