The Full-Stack War: Why Every AI Company Is Building Its Own Chip

The Full-Stack War: Why Every AI Company Is Building Its Own Chip

"The AI war is moving from a model war to a compute sovereignty war." That line from a widely circulated article last week got it backwards. It is not a compute sovereignty war — it is a full-stack war. Sovereignty is the outcome; the full stack is the driver. When your business model depends on someone else's scarcest resource, and that resource rises in price every quarter, you are not running a business — you are working for someone else's chip company.

The footnotes arrived in dense cluster over the past week: Meta announced September mass production of its fourth-generation in-house AI chip Iris, targeting roughly 7 GW of AI compute by 2026 and 14 GW by 2027. DeepSeek was reported to have secretly developed an AI inference chip for a year, recruiting chip-design engineers. Zhipu AI was revealed to be partnering with a domestic ASIC vendor on a dedicated processor. OpenAI released its first custom inference chip Jalapeño, with Broadcom's CEO claiming performance comparable to Blackwell and Google TPU. Anthropic began poaching from OpenAI's chip team. NVIDIA's "super-customer list" is turning into a "competitor list." Why now? Why is every company doing the same thing in the same month? The answer lies in a subtle but fundamental structural shift in the AI industry.

Training Is the Arms Race; Inference Is the Business

For three years, AI companies competed on "whose model is stronger" — a training narrative of bigger parameters, longer runs, more GPUs. In training mode, NVIDIA is irreplaceable: the CUDA ecosystem, NVLink interconnect, and HBM bandwidth would take competitors five years to match. You cannot build an alternative, so you buy. NVIDIA's gross margin holds above 70% and GPU prices only rise.

But from 2025 to 2026, the industry's center of gravity quietly shifted. Training a GPT-5.6 or a Claude Fable happens once a year. Inference — every ChatGPT answer, every agent tool call, every code completion — happens trillions of times daily. AI entered the inference era, and inference math is entirely different from training math: training values peak performance (HBM bandwidth, interconnect speed); inference values marginal cost (how much per token). Morgan Stanley even coined a term — Chip Inflation: GPUs, HBM, optical modules, and switch chips rising across the board, with NVIDIA's H100/B200/Blackwell each generation pricier while AI companies' inference usage grows each generation. It is inflation racing against itself: the more successful you are, the more you depend on GPUs, and the higher your costs climb.

Why ASIC Wins Inference

Training chips and inference chips are two entirely different products. Training demands extreme FP32/FP64 precision, massive HBM capacity, and high-bandwidth interconnect — precisely GPU strengths, since CUDA cores were designed for general parallel computing. Inference's needs are categorically different: cost per token must fall, power matters as datacenter electricity claims a growing share, users cannot wait, and throughput per unit area is the key metric.

ASICs (Application-Specific Integrated Circuits) are born for inference. Google TPU proved it — inference performance and energy efficiency far exceeding same-generation GPUs. Amazon Inferentia proved it again. Now it is OpenAI's Jalapeño, Meta's Iris, and DeepSeek's inference chip's turn. An inference chip does not need 1,000 mm² of silicon or 800W of power. It needs one thing done well: running Transformer inference at minimum cost. The more specialized, the more efficient. Broadcom's custom AI chip business is the best evidence: Q2 semiconductor revenue of $10.8 billion, up 143% year-on-year. Large companies do not need to design from zero — they engage Broadcom to customize, TSMC to produce, and the price-performance beats buying NVIDIA's off-the-shelf product by a wide margin.

The Full-Stack War Doctrine

A clear pattern from the past decade of tech: profit concentrates in companies that control the full stack. Apple is the canonical example — the iPhone captures 85% of the industry's profit not because its chip is the fastest or iOS the best, but because Apple controls silicon (A-series), OS (iOS), the App Store, hardware design, brand, and channels simultaneously. Competitors can win a point; Apple wins the system. AI is tracing the same path. Today's leading AI companies no longer build only models: Google runs Gemini plus TPU plus Cloud plus the world's largest fiber network; Meta runs Llama plus MTIA chips plus 14 GW of datacenters plus the largest social network; OpenAI runs GPT plus Jalapeño plus Stargate datacenters plus the Oracle/SoftBank compute alliance; DeepSeek runs open models plus self-developed inference chips plus domestic chip adaptation; Zhipu runs GLM plus ASIC partnerships plus a domestic cloud ecosystem.

This is not coincidence — it is the full-stack war doctrine: when a technology becomes infrastructure, single-layer competitiveness cannot convert to long-term profit. Profit flows only to companies achieving system-level optimization across chips, models, platforms, and applications. Why? Because the biggest cost-optimization gains lie not in any single layer but in cross-layer coordination. Apple's chip team understands iOS memory management, so they strip unnecessary caches and reallocate silicon area to battery. Likewise, OpenAI's Jalapeño knows it runs Transformer inference, enabling instruction-set-level optimization that slashes per-token cost to a fraction of NVIDIA GPU levels. That kind of cross-layer optimization is impossible with off-the-shelf GPUs.

NVIDIA's Crossroads: Training Holds, Inference Slips

Betting against NVIDIA is dangerous. Every "NVIDIA is finished" narrative of the past five years was proven wrong. H100 sales exceeded all expectations, Blackwell sold out before delivery, and the training market will belong to NVIDIA for years. The real shift is in inference. Inference is the incremental market: as AI moves from demos to daily use, inference will eventually consume a thousand times the compute of training. That is not substitution — it is addition. The question: can NVIDIA capture the increment?

Data has begun answering. Google Cloud increasingly uses TPU rather than GPU for inference, saving 40-60%. Amazon SageMaker defaults to Inferentia for inference. Meta's recommendation and ad-ranking systems have fully migrated to in-house MTIA chips. OpenAI's own inference workloads will gradually shift to Jalapeño. NVIDIA's inference-market share is being carved away by its own customers, slice by slice. Not overnight — but like cloud computing replacing on-premise servers: slow and irreversible. For Chinese AI companies, there is an additional variable: US export restrictions keep the most advanced GPUs out of China. Bernstein projects NVIDIA's share of China's AI chip market falling from roughly 40% in 2025 to about 8% in 2026, with domestic vendors led by Huawei Ascend and Cambricon capturing nearly 80% of China's AI server chip market. In China's full-stack war, chip autonomy is not optional — it is survival.

Three Strategic Positions

AI application-layer founders: do not assume inference costs will keep falling. Short-term, self-chip-wielding giants will cut inference prices faster, squeezing margins for companies without chip capability. Your survival strategy is not competing on price with platforms — it is finding vertical scenarios they do not cover and building moats on product depth rather than cost advantage. Compute middle-layer practitioners: large-scale training remains NVIDIA's territory, but inference infrastructure is fragmenting. That fragmentation is the middle layer's opportunity — multi-chip scheduling, hybrid inference routing, cross-platform cost-optimization engines that the chip-building giants have no incentive to build but customers desperately need. Investors: model scores are becoming a weakening indicator of long-term competitiveness. The better metric: where does a company sit on the inference cost curve? A model with 50% higher per-token costs than competitors cannot hold market share regardless of quality. Conversely, Broadcom — the company selling shovels to shovel-makers — may be the lowest-risk beneficiary of the entire full-stack war.

Sources: AI Business Review, 36Kr, Reuters, The Information, TrendForce, Bernstein.

Scroll to Top