MiniMax Pivots to API Token Sales as Agent Calls Surge

The first post-IPO half-year report from MiniMax, published on August 26, marks a clear business-model pivot: open-platform and API revenue now makes up 63.4% of total revenue, annualized recurring revenue (ARR) has passed $800 million with roughly 80% coming from enterprise customers, and July token consumption hit 20 times the January level. The super-app bet is being shelved — in the agent era, model value is measured in calls per user, not user count.

The super app bet is over

MiniMax was known for two consumer products: Talkie, the AI companion app, and Hailuo AI, the video generator. A year ago, consumer AI-native products contributed nearly 70% of revenue, with the open platform and enterprise services at just 30.3%. In the first half of 2026, those positions almost exactly flipped: platform and other AI enterprise services reached $73.9 million, up 703.1% year over year, or 63.4% of total revenue. Consumer products still grew 100.9% to $42.6 million, but they are no longer the growth engine.

The pivot is not a single-quarter artifact. By August, ARR had surpassed $800 million, with business customers (ToB) accounting for roughly 80% of it. Paid enterprises and developers passed 2 million, about 10 times the level at the end of 2025. And 60.8% of revenue now comes from markets outside mainland China, across more than 230 countries and regions.

Why agents changed the math

The underlying shift is how models get used. In July, MiniMax token consumption reached 20 times the January level, growing far faster than the user base. In the chatbot era, one user request meant a few conversation turns. In the agent era, a single task is decomposed into search, planning, execution, and verification steps, and each step may call the model.

The commercial metric therefore shifts from user count to calls per user. The model cadence at MiniMax tracks this: M3, released in June, sharpened coding and agentic abilities; H3, out in late July, added multimodal generation and was open-sourced.

The cost-per-token race

The token business has different unit economics than SaaS. Every inference consumes real GPU time, electricity, and bandwidth, so revenue growth must be matched by falling cost per token — that is where the margin story lives. First-half gross margin improved from 12.1% to 17.9%, with gross profit at $20.8 million versus $3.7 million a year earlier, attributed to infrastructure efficiency rather than price increases.

The trajectory matters more than the level: text-model compute throughput has roughly tripled in the past two months, and the next-gen M3.1 targets inference cost at about one-third of the M3 launch level. The logic at management level: part of the cost reduction passes through as price cuts to stimulate more token consumption, the rest flows to gross margin — so lower prices and improving margins can coexist.

Open source as a growth lever

H3 has been downloaded more than 24 million times in just over three weeks, spawning more than 300 derivative models. The answer from founder Yan Junjie to whether open source cannibalizes revenue: long-term pricing power depends on capability, cost, stability, and iteration speed — not on whether a model is open-sourced. M3 pricing roughly matches M2 while capability improved, pulling in new customers and pushing existing ones into more scenarios. That makes a cleaner growth chain: better models, more tasks delegated to AI, agents splitting tasks into more calls, higher API revenue.

What it means for the AI industry

The "waiting for a super app" narrative is losing to an "API infrastructure" narrative across China model labs. The recent DeepSeek API price hike shows the same logic from the opposite side: as demand surges, pricing power migrates to whoever controls unit inference cost. The capability race keeps moving too — Rich Sutton argues LLMs need continuous learning to escape local optima, which is exactly the iteration speed MiniMax names as a pillar of pricing power. The competitive axis is now cost per token, reliability, and iteration velocity.

What developers and builders should do

First, design for multi-step agentic workflows and measure cost per completed task rather than per prompt — this is where token economics meets product decisions, together with AI-native engineering practice and durable agent state design. Second, compare providers on unit cost and predictable peak/off-peak pricing, not headline model scores. Third, treat open-source derivatives as a hedge that keeps workloads portable when API pricing shifts. Fourth, watch the actual inference pricing of M3.1 — it will set the unit-cost benchmark other labs must match.

FAQ

Why did MiniMax stop betting on a super app? The numbers moved first: in the first half of 2026, platform and API revenue made up 63.4% of total revenue, up 703.1% year over year, and ARR passed $800 million with about 80% from enterprise customers.

How do agents change how AI companies make money? Agents split one task into search, planning, execution, and verification steps, each of which can call a model. MiniMax July token consumption was 20 times the January level — monetization shifts from user count to calls per user.

Is MiniMax profitable? Not yet. The H1 2026 adjusted net loss was $293 million, though gross margin improved to 17.9% from 12.1%. Management targets M3.1 inference cost at about one-third of the M3 launch level as the path to viable unit economics.

Leave a Comment

Scroll to top