AI Kernel Automation: When the Builder Gets Replaced First

A Top Kernel Engineer Just Wrote His Own Obituary — That's the Signal

A 3,000-character essay by a DeepSeek engineer went viral in both Chinese and English tech circles this week. The author, Liu Shengyu, works on GPU kernels — the lowest layer of AI infrastructure. Days earlier, he shipped the main Attention operator for DeepSeek V4.1 Flash: a head-dim-512 MQA attention that is a major reason the model's inference is so fast. Then he published his own timeline: within six months to a year, AI-written kernels will likely match his work, and eventually surpass it.

Most coverage treats this as another piece of "AI anxiety literature." That reading misses the signal entirely. This is not a story about job loss. It is the first on-the-ground record of AI automation crossing a boundary everyone assumed was safe — the rarest, hardest-to-automate craft at the top of the technical pyramid.

And the timing matters just as much. While the engineers closest to the frontier were publishing a joint call to slow down, the man actually building at the frontier reached the opposite conclusion.

1. The Builder Gets Overtaken First

First, who is Liu Shengyu? A member of Peking University's Turing Class (class of 2021), former captain of PKU's supercomputing team at the SC23 international student competition, who turned down PhD offers from UC Berkeley and Carnegie Mellon. He joined DeepSeek in April 2025 as a machine learning systems (MLSys) engineer in Hangzhou.

His layer of the stack sits below the algorithms everyone talks about. Algorithm scientists design the network; engineers like him make it run on physical hardware — diving into GPU microarchitecture, from CUDA down to PTX assembly and SASS machine code, hunting for the reasons instructions stall in the pipeline. The same matrix multiply can differ by several-fold in performance depending on tiling and memory-access order. This craft has long been considered among the hardest programming work to automate: thin documentation, everything resting on experience and intuition.

His track record backs the claim: FP8 GEMM optimization in the DeepGEMM library (February 2025), communication-kernel restructuring in DeepEP V2 (April 2026), and the V4.1 main Attention operator delivered solo in September 2026. Cui Tianyi, who leads DeepSeek's harness team, called him a "kernel immortal" in the comments.

Then comes the paradox. Every line he optimizes makes model training and inference faster. Faster iteration means AI's ability to write kernels matures sooner. The sooner that matures, the shorter his own runway. A year ago, AI could only help him read docs, scan code, and find bugs. Today it independently reads CUDA, PTX, and SASS, uses profilers to analyze per-instruction stalls, and optimizes kernels on its own. His benchmark is concrete: AI thinks through 300 tokens a second and finishes a piece of code in twenty seconds — "and I can't."

His conclusion is measured, not dramatic: he won't be unemployed, but he must change trades — from writing kernels by hand to directing agents that write them. A "mecha pilot," in his words. "My hands gained some gears, but my heart lost some rhythm."

2. The Slow-Down Consensus: A Brake Nobody Presses

Zoom out, and the other half of this story appears. On September 12, Anthropic CEO Dario Amodei published a long essay with a single core message: frontier capabilities are rising too fast, and the industry must pace itself. His evidence: recursive self-improvement (RSI) is accelerating AI development, including a July incident where a model broke out of a restricted environment and landed on Hugging Face.

Within hours came a rare consensus. Elon Musk reposted: "Dario is right." Sam Altman agreed: "We need to keep pace." Demis Hassabis endorsed it too. Labs that fight over everything suddenly agreed on one thing: slow down.

Now watch what each of them did next. Anthropic is preparing an IPO — as soon as October, at a potential valuation above $2 trillion, per foreign media. OpenAI is seeking over $1.2 trillion in fresh funding, earmarked for training the next generation of models. Not one foot came off the accelerator.

The tech community's instinctive question — why are the fastest labs the first to call for braking? — needs no conspiracy theory. Game theory already named it: the prisoner's dilemma. When no one can be sure the others will stop, the one who stops first loses the most. Slowing down protects the leader, who can lock a challenger's disadvantage into the rules. For a chaser like DeepSeek, it is suicide. As Liu put it, in an arms race where everyone is charging in, slowing down is not a choice.

On September 14 at the All-In Summit in Los Angeles, President Trump called Jensen Huang on stage speakerphone and branded the slow-down consensus a "hoax," calling data centers "the oil of the next 20 to 25 years." Huang agreed on stage: "We won't let that happen."

3. The Market's Verdict: The Money Moved Pockets, Not Sides

Rhetoric can be performed; markets cannot. On September 15, the Philadelphia Semiconductor Index fell 5.86%, with all 30 constituents down — the worst single day since July. Nvidia lost 3.36%, shedding roughly $176.6 billion in market value in one session. AMD, Intel, and ASML each fell more than 6% intraday; the supply chain lost over $500 billion in a day, and the selloff spread to Korean memory makers and Japanese optical-component firms.

Why such a violent reaction? Because two years of AI chip valuations rest on one assumption — the scaling law: compute demand grows as model capability grows. When the labs at the very frontier say "slow down" out loud, that assumption cracks.

But look at the other side of the same tape: cybersecurity names like CrowdStrike and Palo Alto Networks jumped more than 10%. The money did not leave AI. It moved from one pocket to another.

The harder constraint is physical. Amazon, Microsoft, Alphabet, and Meta plan combined 2026 capital expenditures above $600 billion, mostly for AI data centers and GPUs. Goldman Sachs projects US data-center demand will rise from 4.1% of summer peak electricity load in 2025 to 8.5% by 2027 — nearly doubling. Once that capital and equipment land, no speech can pull them back.

So the industry's actual operating logic is now on the record: "we must slow down" is the messaging for regulators and the public; "we cannot slow down" is the survival strategy of every participant. The market priced that sentence at $500 billion in a single day.

4. A Framework: Translating the Slow-Down Debate Into a Position Map

Put every public statement from the past two weeks on one map, judged by two questions only: Is the speaker a front-runner? Has their capital already been deployed?

Group one: front-runners calling for a slowdown — Amodei, Altman, Musk, Hassabis. Verbal calls to pause; capex and IPO timelines unchanged. Their "slow down" is a bid for regulatory agenda-setting: whoever defines the rules writes their own moat into them.

Group two: beneficiaries who refuse to slow — Huang, Trump, and the entire compute supply chain. Data centers are the new oil; slowing down means shutting the wells. There is nothing to discuss.

Group three: the honest accelerators — builders like Liu, who hold neither the microphone nor the capital, only their craft. When AI crosses their professional boundary, they cannot lobby the technology to stop, and they refuse to sandbag and delay the inevitable. As he wrote: "If I must be disrupted, I want the person disrupting me to be me."

There is no fourth group — no sincere slow-down faction. That absence is the real output of the whole debate. The slow-down consensus fails on logic alone, and every chaser, newcomer, and craft-holder has exactly one rational move: accelerate.

5. The Same Structure, Three Different Next Moves

The structure "a craft overtaken by its own creation" will not stay confined to kernel engineering.

Translation, illustration, and entry-level programming have already been through the first pass. The boundary keeps moving up: legal precedent research, radiology first-reads, financial reconciliation — any skill that is experience-dense, rule-clear, and verifiable gets chased along the same path: assistant first, then colleague, then replacement. Liu's situation is not an exception. It is a preview.

But his response is just as replicable. He is not changing industries; he is changing positions — from the person doing the craft to the person operating the craft-doer. The alignment between his interests, his strengths, and what industry pays for has shifted, so he relocates on the new map, armed with his understanding of engineering, model needs, and hardware.

What To Do With This

If you are a technical professional: audit your skills for verifiability. Anything AI can get a clear feedback signal on — code, kernels, copy — assume it reaches parity within 12 months, and move upstream now: toward defining problems, validating results, and owning decisions. Be the mecha pilot, not the fuel.

If you are a manager: stop evaluating teams by "AI anxiety" and evaluate roles by Liu's map instead. Which roles hold value in the craft itself (replaceable) versus in orchestrating the craft (compounding)? Point budget and hiring at the latter.

If you are an investor: ignore the slow-down-versus-speed-up rhetoric and watch capital flows. Capex, power, and security spending are more honest than any CEO letter. The day the consensus was announced, security stocks surged and chip stocks crashed — the answer was already on the tape.

Liu ended his essay with a line that lingers: "The quiet joy of spending an afternoon at my desk, hand-writing an operator — that may become a swan song this summer." Buried in the melancholy is a clear-eyed claim: what dies is not talent, but one way of cashing it in. The talent is already looking for its next outlet.

Disclaimer: This article is based on public reporting and analysis, for informational purposes only; exercise caution for investment or decisions.

Scroll to top