On July 7, Reuters reported that DeepSeek had been secretly developing an AI inference chip for about a year, focused on inference workloads. The same day, Zhipu was revealed to be in preliminary talks with domestic chip-design firms about co-developing a dedicated AI processor.
The immediate reaction was a single question: can a Hangzhou model company build a chip that beats NVIDIA? Domestic processes are stuck above 7nm, US export controls bite — is this not just a press stunt? That question aims in the wrong direction. DeepSeek's real goal in building an inference chip was never to produce something stronger than an H100. Its goal is to ensure that at the negotiating table, NVIDIA can no longer say the one sentence that ends every pricing conversation: "That's the price. Take it or leave it."
Training and Inference Are Different Sports, but Monopoly Pricing Only Cares About One Thing
A technical prerequisite first: training and running a large model place categorically different demands on silicon. Training requires massive parallel gradient computation across billions of parameters, with extreme demands on compute density, memory bandwidth, and interconnect speed. H100s and B200s have genuine dominance here — no alternatives means no alternatives. But inference? Given an input, run one forward pass, output a result. The computation graph is relatively fixed, the generality requirement far lower, and the room for specialization far larger.
Google validated this logic early. TPUv1 launched in 2016 on a 28nm process — while NVIDIA was already at 16nm — yet Google's internal tests showed TPUv1 achieving 15-30× performance improvements and 30-80× energy-efficiency gains over contemporary CPUs and GPUs on inference tasks. A dedicated chip on a mature process beat a general-purpose GPU on an advanced process, because ASIC architecture is tailor-made for one computation pattern. DeepSeek fielding a 14nm or even 7nm inference chip is a completely realistic technical target — categorically easier than "building an H100 competitor."
But technical feasibility was never the core of the story. The core is this: NVIDIA's pricing power rests on the premise "you have no second option," not on the premise "my chip is the best." NVIDIA's gross margin has held above 75% for years. A company sustains that margin on exactly one condition: buyers have no alternative. Training a large model requires H100s. Deploying inference at scale requires GPU clusters. Software ecosystem support exists only inside the CUDA framework. NVIDIA sells not just a chip but an integrated solution whose migration cost is prohibitive. That pricing power is unrelated to cost and entirely a function of alternatives.
The "Second Market Stall" Framework
Picture a wet market. You approach a pork stall; the vendor quotes 28 yuan per jin. You ask for a discount. He says: "That's the price. Take it or leave it — there's no one else." You turn around and see that a second stall has opened across the aisle. The price is not yet set. The pigs are not yet slaughtered. But the door is open. You look back, and the first vendor's tone changes.
That is what DeepSeek's inference chip is for. It does not need to be the strongest chip. It does not need to reach million-unit mass production. It does not even need to complete a successful tape-out. It needs only to exist as a second option at the negotiating table. DeepSeek's compute-procurement lead does not need to say "we're not buying yours anymore." They need only to say: "Our own inference chip enters small-scale tape-out next year. If there is no pricing flexibility, we will prioritize running inference on our own silicon." That sentence is worth every bargaining chip NVIDIA holds. Monopoly pricing has a fragile point: it holds only when buyers believe they have no exit. Even an immature exit — still in tape-out, still yield-limited — is enough to soften the quote. DeepSeek is itself a massive inference-compute consumer: DeepSeek V4 launched in April at 284B-1.6T parameters, and serving it externally consumes inference compute continuously at enormous scale. The chip project is backed by genuine internal demand, not a concept validation. That logic is identical to why Google, AWS, and Microsoft built their own chips: the largest buyers are systematically building backup exits for themselves.
Not Just DeepSeek: the Domino Effect
DeepSeek made headlines, but what makes Wall Street nervous is not one company. Google's TPU program has iterated through eight generations since 2015; TPUv4 on 7nm delivers 5-87% better performance than NVIDIA A100 depending on workload, and Google's internal inference load has shifted heavily to TPU, steadily reducing NVIDIA dependence. AWS's Trainium series is in its third generation, used not just internally — OpenAI and Anthropic have both adopted Trainium 3 for partial workloads. Microsoft is betting on its in-house Maia chip. The largest buyers of AI compute are systematically reducing NVIDIA dependence. They will not stop buying entirely — but where alternatives exist, they are choosing them, leaving NVIDIA for scenarios without options. DeepSeek's entry means even China's largest inference-compute consumer is walking this path. That is a signal of structural demand change.
The gap is real, of course. Google and AWS tape out on TSMC 5nm and 3nm lines; DeepSeek, constrained by sanctions, relies on domestic mature processes — the scale and sophistication are not comparable. But the real tension is that the domino wall just gained another piece. When players of this magnitude are building their own chips, what cloud vendor has grounds not to negotiate with NVIDIA? And DeepSeek's software-stack independence is already substantial: it built hfreduce to replace NVIDIA's NCCL communication library, deployed the 3FS distributed file system, and is already using Huawei Ascend for inference assist. Every move points the same direction: this company does not intend to depend on any external supplier forever.
What It Means
DeepSeek's chip project will not replace NVIDIA. It does not need to. The "second market stall" — even one that is not yet fully stocked — changes the pricing conversation for every buyer in the market. When the largest inference-compute consumers systematically build alternative supply paths, NVIDIA's 75% gross margin faces a structural challenge that no product roadmap can address. The chip itself is almost beside the point. What matters is that the option now exists.
