Everyone assumes the DeepSeek story is about releasing a stronger model. In September 2026, the story quietly changed shape.
In early September, reports emerged that DeepSeek plans to build a data center of roughly 1GW in Ulanqab, Inner Mongolia, deploying at least 160,000 of Huawei's next-generation Ascend 950DT accelerators. At the price leaked on September 10 — over 250,000 yuan per accelerator card — the chips alone represent more than 40 billion yuan in nominal investment. The same week, DeepSeek reportedly hired CITIC Securities to prepare an IPO on the Shanghai STAR Market. And on September 10, DeepSeek released V4.1 Flash, cutting prices again: 0.02 yuan per million tokens for cache-hit input during off-peak hours, 1 yuan for cache-miss input, 4 yuan for output.
Put these three events together and the real story surfaces: the competition is no longer about building a better model. It is about producing massive volumes of tokens at the lowest possible cost. Models are rapidly commoditizing; token prices have fallen to a few yuan per million, while the factories that produce tokens are getting bigger and more expensive. This is not one company's strategy. It is the structural shift of AI's inference era.
Every answer needs a place on the map. This time, it is Ulanqab.
130 Billion Yuan in One Month: Compute Is Becoming an Energy Industry
Start with the numbers. According to IT Times, in August 2026 alone Ulanqab signed 12 intelligent computing projects with a confirmed investment total exceeding 130 billion yuan. A Goldman Sachs report from August 2026 found that Ulanqab's committed data center capacity reached roughly 12.5GW by June — up from 3.3GW a year earlier, nearly a fourfold increase in twelve months.
For scale: OpenAI's Stargate project, backed by a planned $500 billion, targets 10GW. A Chinese city of little more than a million people has, on paper, surpassed it.
The August deal flow deserves a closer look. On August 6, Envision Group's "Star (Gigawatt) River" base came online with roughly 2GW of planned capacity — one of the largest single AI computing sites in the world. In late August, Jining District signed five data center projects in one batch: Sinnet (2GW), Chindata (1GW), ZLUnity (1.5GW), GLP (1.3GW), and Huayi Cloud (500MW) — 6.3GW combined. On August 22, at the China Green Computing Power Conference, another five projects followed, including a gigawatt-scale Volcano Engine facility.
Even more telling is who is investing. IT Times' review of deals since 2024 shows a clear progression: 2024 brought tens-of-billions-yuan data center projects; 2025 introduced "zero-carbon" and "gigawatt-scale" language; by 2026, investors had expanded from IDC operators to cloud providers, telecom carriers, energy companies, and power equipment makers. When computing campuses are measured in gigawatts, they stop being IT projects. They become energy-intensive industry.
The Mengxi Grid: Ulanqab's Real Moat
So why Ulanqab? Guizhou, Gansu, and Ningxia are all building large data centers.
The answer is not just "cold weather." Ulanqab's annual average temperature is about 4.3°C, with nearly ten months of free cooling per year. It sits 320 kilometers from Beijing, connected by two dedicated 144-core fiber routes with one-way latency of about 2.1 milliseconds — making it arguably the closest large-scale compute base to the center of China's AI industry. But these advantages existed ten years ago.
The genuinely scarce asset is the grid. Nearly every Chinese province's grid belongs to State Grid or China Southern Power Grid. Western Inner Mongolia is the exception: it is run by Inner Mongolia Power Group, an independent system called the Mengxi Grid. A historical accident has become an AI-era asset. The Mengxi Grid sits on abundant wind, solar, and coal resources — and runs a long-standing market-based power trading system. In 2024, market-traded renewable energy on the Mengxi Grid exceeded 92% of new energy volume.
The specifics are striking. Documents from the Inner Mongolia Energy Bureau show that by April 2026, two multi-year power purchase agreements for big data enterprises had been filed at an average price of 219.6 yuan per MWh — roughly 0.22 yuan per kWh. Green power trading already covers 83% of the computing industry's electricity consumption in western Inner Mongolia. Some Ulanqab computing centers pay a delivered price of about 0.358 yuan per kWh. By the end of 2025, Ulanqab's installed renewable capacity exceeded 20.3GW, with green power making up 67% of local generation.
The key is not just cheap power. It is predictable power. Multi-year purchase agreements mean a computing campus can forecast — even design — its future energy cost curve. The Zhongjin Data zero-carbon campus, commissioned in July 2025, is the template: instead of buying green power through the grid, it pairs 300MW of wind and solar with 45MW of storage, feeding the data halls directly through a dedicated substation and line. At full capacity it will self-consume 848 million kWh of green power annually, with about 70% of electricity from local renewables. In 2026, Inner Mongolia went further, designating green power parks in Helingeer and Ulanqab as incremental distribution networks where new renewables connect directly to load — the "source-grid-load-storage" model.
A quick calculation shows why this matters so much. A 1GW campus running at high utilization consumes roughly 8.76 billion kWh per year. Save one cent per kWh and you save nearly 90 million yuan a year. Five cents, and the gap exceeds 400 million yuan. For a token factory, the electricity price is the business.
From "Having Compute" to "Having Orders"
But Ulanqab's ledger still shows a glaring deficit.
At MWC Shanghai this year, Zhang Jin, general manager of carrier solutions at Tencent Cloud, offered an uncomfortable number: average GPU utilization at Chinese intelligent computing centers is below 30%. Even the cheapest electricity in China cannot survive that. The China Green Computing Power Conference reported that by August, Ulanqab had signed 109 computing projects totaling over 7 million standard racks — with 650,000 built and 172,000P of capacity actually in operation, against more than 500 billion yuan of cumulative investment, over 95% of it intelligent computing. Meanwhile, an Economic Observer field investigation found land quotas becoming a real constraint, with some signed projects still awaiting construction. A local official admitted the industry is still in its investment phase: "still losing money."
Signing 500 billion yuan is not the same as using it. This is what I call the compute gap: the distance between capacity promised on paper (12.5GW signed) and capacity actually consumed by real workloads (650,000 racks built, 172,000P in operation) spans an entire order chain. At the other end of that chain stands a company like DeepSeek — pricing V4.1 Flash at a few yuan per million tokens while placing an order for 160,000 accelerators on the grassland. Cheap power attracts factories; cheap tokens attract orders; orders fill the GPUs. What Ulanqab must prove is that this flywheel can actually turn.
So How Much Does a Million Tokens Really Cost?
When a million tokens sells for a few yuan, AI's cost war reduces to a single question: how much does it cost to produce them? Half the answer lives in model architecture. The other half lives on the grid dispatch chart.
If you run a model company:
- Put energy costs on the model roadmap. Inference cost equals chip depreciation times electricity price times utilization — all three matter equally. Sign multi-year power purchase agreements to lock in marginal cost.
- Treat utilization as the first KPI. An industry average below 30% means the biggest waste is not on the electricity bill but in idle silicon. Off-peak cache pricing (DeepSeek's 0.02 yuan tier) is how idle capacity turns into revenue.
If you are an investor:
- Separate signed gigawatts from operating petaflops. The first is a promise; the second is revenue. Evaluate compute assets on order coverage and utilization, not headline deal value.
- Watch "green power direct-connect" structures. They may shift the inference cost curve more than the next chip generation will.
If you are a builder: as tokens get cheaper, the cost constraint on inference-heavy applications — long context, agents, video generation — is dissolving. The product ideas you shelved because "we couldn't afford to run it" deserve a fresh calculation.
A token factory on the Mongolian steppe sounds like a joke. But when models commoditize and price wars burn through the floor, what decides who stays at the table is no longer the leaderboard. It is the price of a kilowatt-hour and the utilization of every GPU. That is the second half of AI.
Sources: IT Times (via 36Kr), "Ulanqab, Where DeepSeek Hunts for the Cheapest Tokens"; Reuters reporting; public conference data as of September 2026.
