It's an efficiency-and-revenue kind of morning. MiniMax's annualized revenue jumped 500% on a 2000% spike in token consumption — a clear sign that agents, not just chatbots, are becoming the real engine of AI monetization. Zhipu quietly confirmed its rumored "Niu Lai" model as GLM-5.3 Flash, its first natively multimodal GLM running on domestic chips, while fresh funding keeps flowing into AI infrastructure plays.
Top 3 Highlights
1. MiniMax ARR Surges 500% as Token Use Explodes 2000%
Core insight: The Agent dividend is real: enterprises are burning tokens through agents, and MiniMax is monetizing that shift at a pace few rivals can match.
Source: qbitai.com
2. Zhipu Ships GLM-5.3 Flash: First Natively Multimodal GLM on Domestic Chips
Core insight: The mysterious "Niu Lai" model is confirmed as GLM-5.3 Flash — a fast, natively multimodal GLM running on domestic silicon, a meaningful step for China's homegrown AI stack.
Source: qbitai.com
3. TokenRhythm Raises Tens of Millions, Launches "China's OpenRouter"
Core insight: The AI-infrastructure startup is building an OpenRouter-style model API aggregator and has now stacked up tens of millions of dollars in cumulative funding.
Source: qbitai.com
More News
- Qwen Office debuts Qwen3.8-Flash: the office suite's first release of the new model promises 100% faster generation and 75% lower token consumption. link
- Siemens: industrial agents aren't "wrapped LLMs": a century of engineering know-how is being baked into its industrial AI platform. link
- Silicon Valley's hottest embodied model: learns a task from a single demonstration with no post-training — a possible "GPT moment" for embodied AI. link
- ChronoScale + Microsoft: 50MW AI compute in North America using NVIDIA GB300 NVL72 systems with liquid cooling. link
- S&P: AI infrastructure investment to top $1.3 trillion by 2027. link
Trend Watch
Two threads define this week. First, monetization realism: MiniMax's ARR explosion and iFlytek's pivot from project-based contracts toward platform/MaaS revenue both show the industry has moved from "can we build it?" to "does it make money?" Second, the cost-per-token race: Qwen3.8-Flash cutting token consumption by 75% and Zhipu shipping a Flash-tier multimodal on domestic chips suggest efficiency — not raw scale — is the new competitive frontier.
Disclaimer: This briefing is auto-curated from public RSS sources on August 28, 2026. It does not constitute investment advice.