AI Retrosynthesis Now Matches a 10-Year Chemist

Ten years of bench experience, compressed into five minutes. That is the bar Xia Ning, founder of retrosynthesis startup ZhiHua Technology, now claims for today's AI synthesis design tools — and he has the comparisons to back it up. In internal and external benchmarks, AI retrosynthesis has reached roughly the level of a chemist with a decade of experience: every route a human expert can think of, the model can also propose, and on average it suggests five to six distinct strategies where a person lands on one or two. Clients who spent one to two months hand-designing a route now see it generated in minutes — sometimes with alternatives they had already considered and rejected.

What "10-year chemist" means in retrosynthesis

Retrosynthesis is the planning problem at the heart of drug and material development: given a target molecule, work backwards to find the sequence of reactions that can build it. Xia's benchmark is specific, not marketing fluff. "Basically, any route a chemist can imagine, the AI can imagine too," he says. The edge is breadth — the AI routinely surfaces five to six strategies per target, including routes the client tried and dropped. Clients test this with real cases: they hand over their hardest molecule and compare against a route a senior chemist designed over weeks.

The comparison echoes the 2016 moment when Crystal (now Jingtai Technology) blind-tested polymorph prediction for Pfizer and matched experimental results nobody believed software could produce. Xia says synthesis design has reached the same stage.

How AI designs routes: known reactions, searched at scale

The engine is not generating new chemistry. "Our AI never invents a new reaction," Xia emphasizes. Synthesis innovation almost always reuses known reactions — the trick is that a useful reaction may be buried in a corner of the literature. A chemist who has never seen it will design a 10-step route; the AI, absorbing more papers and retrieving faster, finds the hidden reaction and collapses the route to three or four steps at lower cost. This is verified in production, not aspirational.

Architecturally, the system runs a hybrid of black-box and white-box models. Every conclusion from a large or opaque model must be explained, validated and constrained inside a white-box model before it is trusted. That is the anti-hallucination layer — and, in Xia's view, the invisible commercial threshold for AI in serious science. A black box that occasionally fabricates is not just wrong, it is unfixable: when something breaks, you cannot tell which input caused it. Customers in regulated industries will not buy what they cannot interrogate.

The DMTA loop: synthesis is the bottleneck AI was built for

Drug discovery is not one task but a repeated loop: design, make, test, analyze. The stages run at wildly different speeds. Design is nearly free — thousands of molecules a day. Testing is fast, thanks to high-throughput screening. The make stage is the wall: a chemist synthesizes roughly three to five molecules per month, two to three reactions a day. The entire R&D cycle accelerates only as fast as synthesis does, which is why clients push hardest on this end.

This is a structural fact, not a company pitch: the binding constraint in small-molecule R&D has moved to the exact stage where AI can intervene.

From tool to employee: the Agent-shaped product

Xia describes the post-GPT shift in product form, not just capability. "Previously our technology was presented to chemists as a tool. In the future, the user is an agent, and the UI becomes an API." The vertical algorithms stay intact, but the delivery model changes: a human states the goal, an agent executes the whole workflow, and it keeps a memory of how the work is done — teach it a few times and it behaves like a company employee. That agent-in-the-loop also handles the long tail of generic problems that stumped narrow algorithms: a missing reactor in the lab, an unavailable reagent, the hundred contextual details a specialist system never encoded.

On explainability, the direction is counterintuitive: white-box reasoning is growing, not shrinking, because LLMs explain their decisions. In a serious science setting, if a human remains in the loop, interpretability is non-negotiable — and teams that stay in close client iteration get pushed further toward white-box design.

The token economics of replacing a human chemist

Here is the awkward question investors rarely ask: what does it cost to replace a chemist who earns about 200,000 RMB a year (roughly $28K)? Unlike software, every AI interaction consumes tokens — and that bill decides whether the product gets adopted at all. Xia's team does heavy engineering to keep costs down: route simple problems to cheap or self-hosted models, reserve the best and most expensive models for the critical core, and cut total token spend. Their large clients mostly run self-hosted foundation models, so the cost pressure is manageable today.

Xia's pricing forecast is a hybrid for a long time: SaaS subscriptions for problems the software already solves well, plus per-call agent pricing for problems too hard for SaaS and too expensive for humans. The data philosophy follows a phase model too: positive data matters now (the model learns what worked), negative data becomes decisive later — predicting the reactions people think will work but fail is the only way AI surpasses human judgment.

Why big model labs won't win this vertical

Foundation-model companies are circling AI for Science, but Xia is skeptical they will take the vertical. His internal tests show general models race to 50-60 points on specialist problems and then stall. The gap is data modality: general models are trained mostly on text, while synthesis planning behaves like multimodal data. And a big lab attacking a vertical would have to operate like a startup anyway — without the startup's advantage of a decade of client iteration. "They can use big models, and so can we," he says. "The moat is accumulated understanding of the customer scenario."

What changes for pharma, CROs and investors

  • For R&D teams: benchmark your own hardest cases against AI retrosynthesis now. The gap is not hypothetical — a month of human route design now runs in five minutes, and the AI finds routes your team rejected or missed.
  • For CROs and pharma: synthesis, not design, is where the leverage is. Expect the DMTA loop to compress from the make stage, and expect procurement of chemistry to shift toward agent-facing APIs and per-call pricing.
  • For investors: the market-size question is not "tools" but "full workflow delivered as a result." And the supply-side shift — demand abroad, manufacturing-grade capability at home — could turn China into a "world medicine factory," with India advancing fast too.

Xia's long-term claim is blunt: "Ten years from now, will humans still be doing most of the research? Very likely, AI will be." The era of the AI chemist as a colleague is closer than the industry's caution suggests. For related reading, see how virtual screening benchmarks changed with BoltzMol-1 and how Claude handled autonomous protein design.

Leave a Comment

Scroll to top