The price sheet that ended a price war
DeepSeek changed its API prices twice in seven days. On August 17, the peak-hour output price of its flagship V4-Pro rose from CN¥6 to CN¥27 per million tokens, a 350% increase, while cache-hit input prices jumped 12x from CN¥0.025 to CN¥0.30 per million tokens. Six days later, the company made weekends entirely off-peak. This is not a story about inflation. Read correctly, it is a scarcity report, an announcement that the two-year Chinese LLM price war is over, and the clearest signal yet that the AI compute industry is re-pricing itself — moving from renting GPUs to splitting token revenue with model companies.
Two price moves in seven days
The first change came on August 13, bundled with the release of V4 Pro. From midnight August 17, peak-hour output pricing moved from CN¥6 to CN¥27 per million tokens, and cache-hit input from CN¥0.025 to CN¥0.30. DeepSeek also introduced peak/off-peak (峰谷) pricing: 9:00–12:00 and 14:00–18:00 count as peak; everything else is half price. The second change, effective August 23, made Saturday and Sunday all-day off-peak.
One hike up, one effective cut — they are the same policy. Peak-hour compute is no longer enough for everyone. The invisible hand of price is pushing time-insensitive workloads to nights and weekends, while the scarce daytime slots go to real-time requests willing to pay CN¥27. The market read the first move instantly: on August 17, the A-share optical-module maker Taichenguang hit its 20% daily limit, because investors saw "compute shortage," not "tokens got expensive."
Why the market read it as a capacity report
The demand numbers back that reading. In the week of August 3–9, OpenRouter counted 8.83 trillion tokens of weekly calls on DeepSeek-V4-Flash — up 570% week-over-week and the highest of any model globally. On August 1, the model processed 8 trillion tokens in a single day; by August 4, the API was returning "insufficient capacity" errors. The National Data Bureau puts China's daily token calls at 100 billion in early 2024, 100 trillion by end-2025 and 140 trillion by March 2026 — roughly 1,000x growth in two years.
The shortage is structural, not just viral. Roughly 80% of real-time inference demand sits in eastern China while about 80% of training and batch workloads sit in the west. Data centers can now be delivered in about 100 days, but the power infrastructure takes two to three years. The bottleneck is not only chips — it is electricity and land. Peak/off-peak pricing is therefore scheduling, not promotion. When demand is so strong that capacity must be sliced by the hour, the price sheet itself becomes the evidence of scarcity.
From renting GPUs to splitting token revenue
For two years, Chinese model vendors fought a price war. The logic of a price war is: the product is not yet proven, so give it away first. That logic depressed margins across the compute chain — smarter models, cheaper calls, and upstream orders that felt like one-off trades. The inflection point now shows up in prices. Morgan Stanley's August 9 report, titled "Farewell to Price War, Enter the Intelligence War," tracked official quotes from eight Chinese vendors (ByteDance, Alibaba, Baidu, Tencent, MiniMax, Zhipu, Moonshot AI, DeepSeek): average API input prices were CN¥4.9 per million tokens in Q2 2026 and output CN¥21.9, up 48% and 80% respectively from Q1 2025. Brokers cite Zhipu raising prices three times, Tencent Cloud twice, with Alibaba and Baidu clouds following. The company that started the price war by slashing token prices is the one now raising them 350% — a strong sign the war is over.
Something deeper is shifting in how compute is sold. Traditional compute services rent GPUs — fixed fees per card per hour. On July 29, Xingyun Technology disclosed an upgraded five-year contract with a major model customer (widely believed to be Moonshot AI): contract value raised from CN¥1.014 billion to CN¥3.053 billion, compute units expanded from 128 to 256, monthly fees on delivered units up 21.21% and on new units up 36.36%. The reported new clause is a "token-revenue-linked fixed service fee" — reportedly the first time token revenue sharing has been written into a major A-share contract. Renting hardware makes compute a cost; sharing token revenue makes it an asset. Citic Securities frames it as leasing moving from fixed monthly rent to actual token-usage billing; Galaxy Securities puts it bluntly: from selling resources to selling output.
The chain re-prices itself — and hits two constraints
Money is flowing up the chain, and every link just reported earnings. Optical-module leader Zhongji Innolight posted H1 revenue of CN¥41.78 billion, up 182% year-over-year, with net profit of CN¥13.65 billion, up 242%; management says orders that used to roll on three-month cycles are now being signed into 2027. Wafer foundry Hua Hong recorded a record Q2 revenue of $717.5 million at 102.8% utilization, attributing ~60% of growth to price increases and ~40% to capacity expansion. Equipment maker AMEC saw H1 net profit grow more than 300%. TrendForce raised its 2026 global AI-server shipment growth forecast from 28% to about 31%, and expects the nine major cloud providers to grow capex ~90% — Microsoft, Amazon, Alphabet and Meta alone guiding $735–760 billion in 2026 capex combined.
On August 24, the same day the market sold off optical modules (CSI 300 index down, Zhongji Innolight -7.4% on the A-share and -12% in Hong Kong, with main-fund net outflows exceeding CN¥3.9 billion), Alibaba completed a placement of 710 million new shares at HK$112.70 each, raising HK$80 billion for AI infrastructure — oversubscribed within an hour at nearly 3x, with sovereign and long-term funds taking more than 40%. Two pockets of money made opposite choices on the same day: secondary-market traders selling hardware names, long-term capital queuing to fund AI infrastructure. The market has moved from buying grand narratives to buying evidence.
Two constraints matter. First, price elasticity: around the hike, Alibaba open-sourced Qwen3.8-Max (2.4 trillion parameters), the first Max-level flagship in the open pool — every yuan of premium in the price hike lowers the switching cost for competitors' alternatives. Second, policy risk: Zhongji Innolight derives ~94.8% of revenue overseas, and on August 4 the US FCC was reported to be drafting import restrictions on 800G/1.6T optical modules. Any movement in technology route or trade policy can re-rate the chain regardless of signed orders.
What to watch and what to do
Three indicators decide whether the hike sticks: whether call volumes hold, whether off-peak utilization rises, and whether the cache-hit price gap narrows. If users stay and batch traffic migrates to nights and weekends, the price sheet becomes the industry's thermostat — measuring both demand heat and route sentiment.
For developers and platform teams: move batch and offline workloads to off-peak windows (half price), design for cache hits (input at CN¥0.30 instead of full price), and instrument cost per token rather than gross spend. For compute buyers: evaluate token-revenue-linked contract structures instead of fixed GPU rent. For investors tracking the debate around token taxation (see Gates's proposal to tax tokens), the chain's re-pricing is now measurable — and how token economics are priced on the basis of five layers of cost from silicon to service determines who captures the margin. Those who wait to see whether the Qwen open-source flagship forces prices down again are waiting on the same three indicators.
FAQ
Why did DeepSeek raise API prices by 350%?
Because demand outgrew capacity. OpenRouter counted 8.83 trillion weekly tokens on V4-Flash in early August (+570% week-over-week), and DeepSeek's API began returning capacity errors by August 4. The price move is a scheduling tool: peak/off-peak pricing pushes batch workloads to nights and weekends while scarce daytime slots go to real-time requests.
Is the Chinese LLM price war actually over?
The data says yes. Morgan Stanley's August 9 report showed average API input prices at CN¥4.9/million tokens and output at CN¥21.9 across eight Chinese vendors in Q2 2026, up 48% and 80% from Q1 2025. Zhipu raised prices three times, Tencent Cloud twice, and Alibaba and Baidu clouds followed.
How should developers respond to peak/off-peak token pricing?
Move non-real-time workloads to off-peak hours and weekends at half price, maximize cache-hit ratios (CN¥0.30 vs full input price), and watch the three indicators — call volume, off-peak utilization and cache price gap — to see if the increases hold.