AI Daily Briefing – 2026-08-28: MiniMax ARR Soars 500%

It's an efficiency-and-revenue kind of morning. MiniMax's annualized revenue jumped 500% on a 2000% spike in token consumption — a clear sign that agents, not just chatbots, are becoming the real engine of AI monetization. Zhipu quietly confirmed its rumored "Niu Lai" model as GLM-5.3 Flash, its first natively multimodal GLM running on domestic chips, while fresh funding keeps flowing into AI infrastructure plays.

Top 3 Highlights

1. MiniMax ARR Surges 500% as Token Use Explodes 2000%

Core insight: The Agent dividend is real: enterprises are burning tokens through agents, and MiniMax is monetizing that shift at a pace few rivals can match.

Source: qbitai.com

2. Zhipu Ships GLM-5.3 Flash: First Natively Multimodal GLM on Domestic Chips

Core insight: The mysterious "Niu Lai" model is confirmed as GLM-5.3 Flash — a fast, natively multimodal GLM running on domestic silicon, a meaningful step for China's homegrown AI stack.

Source: qbitai.com

3. TokenRhythm Raises Tens of Millions, Launches "China's OpenRouter"

Core insight: The AI-infrastructure startup is building an OpenRouter-style model API aggregator and has now stacked up tens of millions of dollars in cumulative funding.

Source: qbitai.com

More News

  • Qwen Office debuts Qwen3.8-Flash: the office suite's first release of the new model promises 100% faster generation and 75% lower token consumption. link
  • Siemens: industrial agents aren't "wrapped LLMs": a century of engineering know-how is being baked into its industrial AI platform. link
  • Silicon Valley's hottest embodied model: learns a task from a single demonstration with no post-training — a possible "GPT moment" for embodied AI. link
  • ChronoScale + Microsoft: 50MW AI compute in North America using NVIDIA GB300 NVL72 systems with liquid cooling. link
  • S&P: AI infrastructure investment to top $1.3 trillion by 2027. link

Trend Watch

Two threads define this week. First, monetization realism: MiniMax's ARR explosion and iFlytek's pivot from project-based contracts toward platform/MaaS revenue both show the industry has moved from "can we build it?" to "does it make money?" Second, the cost-per-token race: Qwen3.8-Flash cutting token consumption by 75% and Zhipu shipping a Flash-tier multimodal on domestic chips suggest efficiency — not raw scale — is the new competitive frontier.

Disclaimer: This briefing is auto-curated from public RSS sources on August 28, 2026. It does not constitute investment advice.

Leave a Comment

Scroll to top