AI Daily Briefing – 2026-08-15: Qwen3.8 Goes Open Source

Saturday, August 15, 2026. A big open-source day: Alibaba released Qwen3.8 — including a 27B model that runs on consumer GPUs and a 2.4T-parameter flagship that puts Qwen-Max-class capability in the open. OpenAI published a deep engineering look at GPT-Live's real-time voice system, and Zhipu shipped GLM-5.3 with a security edge. Here's what matters.

Top 3

1. Qwen3.8 goes open source — all the way from 27B to 2.4T

Alibaba released Qwen3.8-27B, free to download, deploy, and use commercially — sized to run on a consumer GPU. On the other end of the scale, Qwen3.8-2.4T-A95B is the largest, strongest open model the Qwen series has ever shipped: 2.4T total parameters, 95B active, with native 262K context extendable to 1M. On Terminal Bench 2.1 it jumped from 74.5 to 86.6, and on SWE-bench Pro it now sits near Opus 4.8. FlagOS had it Day-0 adapted to nine chip platforms including NVIDIA, Huawei Ascend, and Moore Threads, with INT8 quantization paths for existing datacenter fleets. [QbitAI] [Leiphone]

2. Inside GPT-Live: how OpenAI made 95% of audio frames stop lagging

OpenAI's engineering post on its continuous voice system reveals a full-stack rewrite: the p95 audio-frame latency of the new media stack now matches the old system's p50, a custom WARP protocol cuts WebRTC channel setup from 6 network round-trips to 1, and the media frontend moved from Python asyncio to Go. Search, tool calls, and deep reasoning were moved off the audio path entirely. The takeaway: real-time voice is now a systems problem — scheduling, state, and infrastructure — not just a model problem. [Leiphone]

3. GLM-5.3 ships: coding closes in on Fable 5, security leads the open field

Zhipu released GLM-5.3, with coding performance reportedly approaching Fable 5 and a strong showing as the best open-source security model — it even surfaced a bug that had lurked in a codebase for 40 years. The release keeps up the cadence of Chinese labs shipping frontier-adjacent open models within days of each other. [QbitAI]

More news

  • Google launches Gemini 3.7 Flash. The new Flash is better at debugging and can generate near-production-ready code in a single pass, with lower token costs; Gemini Spark now runs on it. [36Kr]
  • NVIDIA's CPO switches enter mass production. Spectrum-X Ethernet Photonics, built for gigawatt-scale AI factories, cuts laser count 4x and power draw 5x versus conventional optics, with TSMC, Lumentum, and Foxconn in the supply chain. [36Kr]
  • Kimi K3, under the hood. Su Jianlin's technical review of K3 details 896 routed experts (16 active) over 2.8T parameters, LatentMoE's compressed expert space, and a KDA + Gated MLA hybrid attention — a blueprint for scaling capacity without scaling compute per token. [Leiphone]
  • Tencent's WeChat AI is staffing up. Hunyuan senior researcher Xu Can (creator of WizardLM) moved to the WeLM team, deepening WeChat's self-developed model and Agent push around the "Xiaowei" assistant. [36Kr]
  • Moonshot AI's IPO timing tightens. Reports say it may file for a Hong Kong listing before September 30; with Kimi K3 ranked third on Artificial Analysis, the window to convert technical buzz into valuation is now, as DeepSeek V4 Pro undercuts on cost. [36Kr]
  • China's AI hardware boom is showing up in the numbers. Huaqiangbei AI glasses sales are up 100% in the first seven months, and embodied-AI funding hit ¥93.5B in H1 2026 — up 5x year over year. [36Kr] [36Kr]
  • Claude chat links reportedly showing up in Google search. Leiphone reports that some Claude conversation shares are being indexed and discoverable via search — a reminder to treat share links as public. [Leiphone]

Trend watch

Three shifts worth tracking. First, open source is scaling up, not just down: Qwen3.8 spans a consumer-GPU 27B at one end and a Qwen-Max-class 2.4T model at the other — open weights now cover the whole capability spectrum, and Day-0 multi-chip adaptation is becoming table stakes for big releases. Second, real-time voice is the new frontier in engineering: GPT-Live's teardown shows the differentiator is no longer the model alone but latency, state management, and systems design around it. Third, the price war is spreading: DeepSeek V4 Pro's stable pricing and US labs cutting mid-tier prices by roughly a quarter in a month mean cost, not capability, is increasingly the binding constraint on who can build with frontier models.

Disclaimer: This briefing is compiled by AI from public sources; verify details at the linked originals. Prices and facts are as reported at publication time.

Leave a Comment

Scroll to top