When Meituan released LongCat-2.0, the tech community's narrative split rapidly into two camps. One hailed it as a milestone of "domestic compute overtaking on the curve" — 1.6 trillion parameters, zero NVIDIA, a 50,000-card Ascend cluster, every word hitting emotional targets. The other calmly noted that its SWE-bench Pro score of 59.5 trailed Zhipu's GLM-5.2, and its anonymous performance on OpenRouter was merely "a lightly modified reproduction of DeepSeek-V4-Pro." Both camps were right. Neither hit the point.
LongCat-2.0's real significance lies not in whether it "surpassed OpenAI" or "overtook on a curve." Its true value is that it simultaneously emitted two structural signals: first, the MoE-architecture trillion-parameter arms race has hit its ceiling; second, China's AI industry for the first time completed a full-chain industrial-grade closed loop — from pretraining through inference deployment — without depending on a single NVIDIA GPU. The former concerns the endpoint of the technical route; the latter concerns supply-chain restructuring under geopolitical pressure. Two signals stacked together define the new coordinates of AI competition in mid-2026.
"Zero NVIDIA": an Overhyped Narrative, an Underrated Engineering Achievement
First, unpack the most eye-catching label: "the world's first trillion-parameter model with zero NVIDIA content." It is overhyped because, looking purely at model capability, LongCat-2.0 did not rewrite any technical paradigm. 1.6T total parameters / 48B active, 1M context window, MoE with sparse attention — this formula sits in nearly the same range as DeepSeek-V4-Pro (1.6T/49B/1M). A highly-upvoted Reddit r/LocalLLaMA comment said it directly: "All those parameters to not be SOTA." All those parameters, and it is still not the strongest.
But it is underrated because most people conflate two entirely different engineering difficulties: "thousand-card post-training" and "50,000-card from-scratch pretraining." The early-June Shenzhen-Huqiao Institute collaboration with HIT and Huawei — completing DeepSeek-V4-Pro full-parameter post-training on Ascend 910C — solved "can domestic chips fine-tune a trillion-parameter model." Meituan did something else: on roughly 50,000 domestic chips, starting from a randomly initialized 1.6-trillion-parameter network, it pushed the model to convergence with over 30T tokens of pretraining data. Any single loss spike, any NCCL communication timeout, any silent data corruption (SDC) on one card could have vaporized tens of millions of dollars in electricity costs. Meituan's disclosed engineering metrics: training MFU improved from 17.8% to 27.68% (1.5×), key operator efficiency up 14%, daily failure rate dropping from 15.7 to 4.4 per 10,000. Supporting these numbers is a complete suite of domestic-chip operator rewrites, parallel-scheme restructuring, and automatic fault recovery — from anomaly detection to link failover to auto-recovery, nearly all automated.
So "zero NVIDIA" is accurately read not as "Meituan beats OpenAI" but as "China's AI supply-chain autonomy has advanced from 'can it work' to 'is it stable.'" This is an infrastructure-level breakthrough, not a model-level one.
97% Sparsity: Why the Trillion-Parameter Arms Race Is Already Over
A sentence in LongCat-2.0's official blog matters more than the "1.6 trillion" number: the model's MoE sparsity has reached approximately 97%; adding another 135B in expert parameters yields negligible performance gains. That sentence deserves to be framed. Trace the MoE evolution: DeepSeek-V3 was 671B total / 37B active, sparsity ~94%; V4-Pro was 1.6T/49B, sparsity ~97%; LongCat-2.0 also 1.6T/48B, sparsity 97%. Three model generations, from 560B to 671B to 1.6T, sparsity climbing from 94% to 97% — then stopping. This means the marginal return on adding more total parameters (while keeping active parameters flat) has effectively reached zero. The industry has hit a scaling wall — not the scaling-law wall of "more data = better" but the architectural wall of "more experts = negligible improvement." The trillion-parameter race is over not because anyone declared it finished but because the math says further parameter scaling without proportional activation increase is wasted silicon.
The implication for the industry: the next competitive frontier is not total parameters but what you do with the active ones — better routing, better expert specialization, better inference efficiency. LongCat-2.0's contribution is not that it won the parameter race but that it proved the race has a finish line, and we are standing on it.
Two Signals, One Coordinate
Stack the two findings — the engineering-maturity signal and the architecture-ceiling signal — and a coordinate emerges for mid-2026 AI competition. The hardware-software gap between China and the US is narrowing not because Chinese chips are catching up to NVIDIA (they are not, yet) but because Chinese teams have proven they can run the entire AI pipeline — from random initialization to serving production traffic — on non-NVIDIA hardware. Meanwhile, the architecture ceiling means the next breakthrough will not come from stacking more parameters but from fundamentally different approaches: better training objectives, better data curation, better inference-time compute allocation. The teams that recognize both signals — infrastructure independence and architectural saturation — will be the ones that navigate the post-scaling era successfully. The ones still racing to 2T parameters are running a race that has already been won.
The "zero NVIDIA" label also carries a geopolitical weight that model benchmarks cannot capture. US export controls were designed to prevent China from training frontier models. LongCat-2.0 is proof that the controls have failed — not because Chinese chips match NVIDIA (they do not) but because Chinese engineering teams have built the systems-level software that makes adequate chips sufficient for frontier-scale training. The bottleneck has shifted from hardware access to software engineering capability — and that is a bottleneck that export controls cannot address, because it lives in the codebases of companies like Meituan, DeepSeek, and Alibaba rather than in any single chip or fab.
