Ulanqab and DeepSeek: The Energy Arbitrage Behind China's Cheapest Tokens

A Model Company Went to the Grassland to Buy Electricity

DeepSeek made headlines again this month — but the real story isn't a model. It's a city most people couldn't find on a map: Ulanqab, Inner Mongolia.

The popular framing is "DeepSeek is building a data center in Ulanqab," as if it were just another site-selection announcement. Line up the recent news, though, and this is not site selection. It is a cost manifesto. In early September, reports emerged that DeepSeek plans to build a roughly 1 GW data center in Ulanqab, deploying at least 160,000 of Huawei's next-generation Ascend 950DT AI accelerators. At the reported price of over 250,000 yuan per accelerator card, the chips alone represent a nominal investment above 40 billion yuan. Back in April, DeepSeek posted a job listing for a senior data center operations engineer — up to 30,000 yuan a month, located in Ulanqab. On September 9, Reuters reported that DeepSeek had hired CITIC Securities to prepare for an IPO on the Shanghai Stock Exchange's STAR Market. On September 10, it shipped DeepSeek V4.1 Flash, with cache-hit input pricing during off-peak hours dropping to 0.02 yuan per million tokens.

Put the IPO, the chips, the electricity prices, and the model pricing on one chart and the conclusion is hard to escape: the second half of the AI race is shifting from "whose model is smarter" to "whose cost of producing tokens is lower." Ulanqab is not a backdrop. It is the solution to that cost equation.

Tokens Get Cheaper; Token Factories Get More Expensive

Start with the demand side. V4.1 Flash's pricing is a mood ring for the industry: 0.02 yuan per million cache-hit input tokens off-peak, 1 yuan on cache miss, 4 yuan for output, doubled at peak. Inference prices keep falling, with no floor in sight. When a million tokens sells for a few yuan, the model itself can no longer carry the differentiation. What decides winners is how many cheap tokens each kilowatt-hour produces after passing through a GPU.

Now the supply side. Token factories are becoming astonishingly capital-intensive. The rumored Ulanqab project would sink over 40 billion yuan into chips alone. Envision Group's Ulanqab Star (Gigawatt) River base, which began production on August 6, is planned at roughly 2 GW — billed as one of the world's largest single AI computing campuses. A Goldman Sachs report on China's data center industry released in August 2026 sets the macro marker: as of June, Ulanqab's committed data center capacity totaled about 12.5 GW. A year earlier, that figure was 3.3 GW. For comparison, OpenAI's $500 billion Stargate project targets 10 GW.

Capital expenditure is sliding from the algorithm layer down to the infrastructure layer. That is not one company's choice; it is a structural migration. When inference becomes AI's primary workload, the word "compute" regains its physical meaning: it is, in essence, a converter that turns electricity into tokens. The question becomes: where in China is that conversion cheapest?

Energy Arbitrage: Ulanqab's Three-Layer Ledger

Western China has been pitching data centers for a decade — Guizhou, Gansu, and Ningxia all have their projects. Why are gigawatt-scale projects converging on Ulanqab right now? The answer is not "it's cold there." It is a three-layer arbitrage structure.

Layer one: geography. Ulanqab sits about 320 km from Beijing, an hour and a half by high-speed rail. Two 144-core point-to-point dual-route fiber cables run straight to the capital, with measured one-way latency as low as 2.1 milliseconds. For inference workloads, that makes Ulanqab the closest large-scale compute base to China's AI industry center — not a backup site stranded in the northwest, but Beijing's compute suburb. Add a climate of roughly 4.3°C average annual temperature and nearly ten months of free cooling per year, and a fixed cost line item is slashed before you build anything.

Layer two: grid institutions. Inner Mongolia's grid structure is unique in China: the east is part of the State Grid system, but the west is run by Inner Mongolia Power Group as the independent Mengxi (Western Inner Mongolia) grid — and Ulanqab sits inside it. This relatively independent, market-oriented system produces numbers few other regions can match: in 2024, market-traded renewable electricity on the Mengxi grid exceeded 92%. According to a document from the regional energy bureau, by the end of April 2026 Inner Mongolia had registered two multi-year power purchase agreements for big-data enterprises at an average price of 219.6 yuan per MWh (about 0.2196 yuan per kWh); green power accounted for 83% of consumption by the compute industry in western Inner Mongolia; and some Ulanqab compute centers pay a delivered price of about 0.358 yuan per kWh. The mechanism matters more than any single number: multi-year PPAs give compute operators a predictable forward energy cost curve — they can design their own energy structure. By the end of 2025, Ulanqab's installed renewable capacity surpassed 20.288 GW, with green power making up 67% of the city's electricity mix.

Layer three: physical direct connection. China JinData's zero-carbon computing base in Ulanqab, commissioned in July 2025, is the template: rather than buying green electricity through the grid, the campus pairs 300 MW of wind and solar plus 45 MW of storage, feeding the data center directly through a dedicated substation and lines. At full capacity it expects to self-consume 848 million kWh of green power annually, with roughly 70% of electricity coming from local renewables. In 2026, Inner Mongolia pushed the model further with an "incremental distribution network + direct green power supply" scheme, designating green-power parks like Horinger and Ulanqab's Chahar as incremental distribution networks where new renewables connect locally and are consumed locally. Electricity no longer detours through the public grid for transmission and settlement; it flows from turbine to GPU.

One back-of-envelope calculation shows why this structure justifies a 40-billion-yuan bet: a 1 GW load running at high utilization consumes roughly 8.76 billion kWh a year. Save one fen (0.01 yuan) per kWh and you save nearly 90 million yuan annually; five fen, and the gap exceeds 400 million yuan. Energy arbitrage is not a metaphor. It is a real line item in the financial model.

The Bottleneck Is Not Electricity. It Is Utilization.

Cheap electricity solves inputs, not production-line efficiency. At MWC Shanghai this year, Zhang Jin, general manager of Tencent Cloud's carrier solutions, offered an unflattering number: average GPU utilization at Chinese AI data centers is below 30%. In other words, even if Ulanqab offers the cheapest power in the country, a utilization gap can swallow the entire electricity advantage.

The signing numbers are spectacular: as of August, Ulanqab had signed 109 computing projects totaling over 7 million standard racks (650,000 built), 172,000 P of deployed compute, and over 500 billion yuan in total investment, more than 95% of it intelligent computing. But an Economic Observer field investigation found land quotas becoming a real constraint, with some signed projects still waiting to break ground; a local official admitted the compute industry is "still losing money." In August alone, Ulanqab signed 12 intelligent-computing-related projects worth over 130 billion yuan in confirmed investment — VNET at 2 GW, China Unicom Cloud at 1.5 GW, GLP at 1.3 GW, Chindata at 1 GW. Supply-side enthusiasm is abundant. Demand-side confirmation is not.

This is the final piece the energy-arbitrage framework has to account for: arbitrage only works when the asset is fully used. A 1 GW data center without sustained high-load inference orders is simply the most efficiently powered idle machine in the country. Ulanqab's real exam question is not "does it have compute" but "does it have orders" — the leap from a lowland of production inputs to a highland of token production-line efficiency.

The Framework Travels — and What To Do

Apply "energy arbitrage + utilization" elsewhere and it holds: US data centers migrating to Texas and other low-price, fast-interconnect states follow the same logic; cloud providers scheduling inference workloads toward renewable-rich regions is the software version of the same problem. For any business that converts electricity into tokens, the first term of the site-selection function is shifting from "close to talent" to "close to cheap power and close to customers."

Concrete takeaways:

  • If you run a model or application company: decompose your inference cost. The marginal return on architecture optimization is shrinking; power purchase agreements and site selection may offer more. Negotiate electricity the way you negotiate funding.
  • If you invest in compute infrastructure: ignore headline gigawatts. Watch two numbers — the signed-to-groundbreaking conversion rate and actual GPU utilization. With the industry averaging under 30%, most projects are losing money on paper by construction.
  • If you are a local government in western China: land and power subsidies buy ribbon-cutting ceremonies. What retains orders is latency, grid reliability, and green compliance — overseas customers' requirements for renewable share will only harden. Ulanqab's Mengxi grid institutional advantage is precisely the part that is hardest to copy.

When a million tokens sells for a few yuan, the AI cost war boils down to one question: what does a million tokens really cost? Ulanqab's answer: ask the grid.

Sources: IT Times (via 36Kr), Reuters, Goldman Sachs "China Data Center Industry Report" (Aug 2026), and public reporting.

Scroll to top