Everyone wants to know who is "winning" the AI race. That question is usually answered with a leaderboard of models. The analyst Qin (董洁林) argues this is the wrong scoreboard — the real game is played across three separate ledgers: the company, the nation, and the individual. A US firm can post record revenue while hollowing out white-collar jobs; Beijing can claim industrial scale while ordinary workers see none of the productivity dividend. The two superpowers will both declare victory, but what victory means for each is structurally different.
A six-layer stack, and where the lock-in actually lives
Beyond Huang's five layers (power, chips, infrastructure, models, applications), the analyst adds a sixth: the Harness layer — the software glue that coordinates models and agents. The image matters: if a model is a horse, the Harness is the reins and saddle that decide how it works. Technical frontier and value capture do not sit in the same layer. A superior model sets the ceiling on intelligence supply, but durable customer lock-in happens upstream, in the apps and orchestration that own client data, workflows, and distribution. This is where Silicon Valley startups are thinnest in assets and most active: converting general intelligence into a specific business outcome.
China won distribution, not the value chain
Chinese open-weight models — Qwen, DeepSeek, Moonshot, Zhipu — have racked up real global traction (see the deeper dive on how Chinese open-weight models power American AI). Hugging Face says Chinese-source models now account for 41% of downloads; DeepSeek is the #1 supplier on OpenRouter by token share. But this is distribution, not revenue. Vercel data shows Chinese open models drive about a third of its gateway's token usage yet less than 4% of its billing. Downloads are not deployments, deployments are not sustained usage, and usage is not income. Open weights buy reputation and ecosystem entry while eroding direct monetization — a large share of the value ends up in the American chip, cloud, routing, and Harness providers underneath, part of the broader US-China infrastructure race. China is winning the top of the funnel, not the ledger.
Token prices collapsed. That rewrites the rules of competition.
The industry's most consequential change of the past three years is not smarter models but cheaper equivalent intelligence: token prices for comparable capability have fallen more than two orders of magnitude. Falling faster than Moore's law can explain, the drop is an ecosystem outcome — quantization, speculative decoding, caching, routing, better utilization — not a single layer's achievement. In late July 2026 OpenAI pushed GPT-5.6 Luna to free users and cut API pricing to $0.20/M input and $1.20/M output. Chinese vendors, meanwhile, are raising prices as subsidies retreat and full-stack TCO finally shows up on the price tag. Cheap tokens are not cheap compute: a head-to-head TCO on effective work still favors NVIDIA stacks in many cases, offsetting China's advantages in engineering and labor.
Three falsifiable predictions
The essay ends with three scenarios, each with explicit falsification criteria — the useful kind of forecast.
1. The US open-weights camp expands. Meta releasing MuseSpark 1.2 weights on Aug 10 is the signal. Expect more top-tier American models to open up, lifting US open weights' share of routed tokens. Open will be selective, though: the most dangerous frontier capabilities stay behind controlled closed APIs. Falsified if US open models fail to gain token share in 18 months or Chinese models clearly dent US model revenues.
2. The model layer consolidates after two brutal years. Expect escalating competition through roughly 2028, then shutdowns and M&A among vendors without stable revenue or unique capability — unless a decisive AGI breakthrough re-orders the field around the leader. Falsified if no verifiable model-layer M&A appears in either ecosystem after that point.
3. The AI digital border approximates today's internet geography. US-platform-dominant zones, a Chinese-independent sphere, and hybrid regions. The line is set not by open competition but by state policy: chips, clouds, payments, app stores, data rules, procurement, and safety standards — the same forces that are already forcing 35 nations to pick a side. Short-term US tolerance of Chinese models is conditional — any visible security or political incident can flip tolerance into a ban. Falsified if a non-aligned third pole (Europe/India) emerges or the Chinese stack dominates multiple big economies outside its historic reach.
What to do with this
For operators and founders, the takeaway is to anchor defensibility in the Harness and application layer — the data, workflows, and distribution that survive model swaps. Favor open weights to reduce supplier dependence, but do not assume cheap models mean cheap compute once TCO is counted. For individuals, watch where the marginal productivity gain actually lands: is your AI use expanding your judgment and client relationships, or concentrating the rent upward? The question to keep asking is whose scoreboard you are winning on.
