Claude's Autonomous Protein Design Hits 14 of 15 Targets

Anthropic just proved that a general-purpose model can run an entire protein design campaign on its own — no computational biology specialist in the loop. Given nothing but 15 target names, Claude designed new proteins that bind 14 of them, and its top designs latch onto targets several times more tightly than the best previously published result.

The task: designing proteins, not just predicting them

AlphaFold-style models answer one question: given a protein sequence, what shape does it fold into? The input is known; the output is a prediction. What Claude did is a different problem — de novo design. The input is only the name of a target protein; the output is an entirely new protein engineered to bind it. One is captioning a picture; the other is writing an original essay.

This matters because binding is how a large share of modern medicines work: a molecule attaches to a target to inhibit, activate, or deliver something. Historically, designing a new binder took protein engineers months of computation, optimization, and screening per target — even after ML models sped up parts of the pipeline, orchestration still demanded a computational expert for days or weeks.

How Claude ran the campaign: orchestration, not invention

Claude invented no new tools. It used off-the-shelf open-source software — RFdiffusion, ProteinMPNN, ESMFold2 — that any lab can download. What Claude contributed was orchestration.

Anthropic encoded roughly 16,000 words of protein-engineering know-how into a prompt: the stages of an experiment, the tools available at each stage, and the screening criteria. The prompt deliberately did not specify which face of the protein to attack, which generation method to use, or any starting sequence. Given a cloud account and network access, Claude picked targets, chose epitopes, loaded tools, ran models, filtered and optimized results, ranked them, and handed off a deliverable — ultimately invoking 10 structure-generation methods and composing 24 tool combinations. Humans did exactly three things: approve network access, monitor infrastructure, and send the ranked designs to two independent wet labs for synthesis and testing.

The results: 22-35% hit rates versus a 10-15% baseline

Adaptyv Bio and Twist Bioscience independently built and tested Claude's designs. Against 15 targets, Claude produced working binders for 14. Its per-design success rate ran 22-35% depending on the setup — two to three times the 10-15% typical of protein design campaigns today. Looking only at Claude's top-ranked design per target, the hit rate reached 49%.

Two targets deserve special mention. On RBX1, Adaptyv Bio had previously run a public design competition: 245 submissions from global teams, only 9 worked. Claude submitted 90 designs on the same target; 28 bound successfully, and its best design bound roughly ten times more tightly than the competition winner. On TNFα — the target behind Humira, one of the world's best-selling drugs — multiple expert teams had failed at de novo binder design. Opus 4.8 produced 12 valid designs, some binding across human, monkey, and mouse TNFα simultaneously. Not everything worked: on one target (MBP), all 90 designs failed — a useful reminder that the pipeline is fast, not infallible.

The chemistry add-on: 23 minutes versus an hour by hand

The second experiment moves from biology to analytical chemistry. Given raw NMR and LC-MS files from a contract lab and a two-sentence prompt — no training, no specialist setup — Claude Opus 5 returned finished analyses in 19-23 minutes. Its hydrogen counts and purity readings (96.4%) matched the lab's own manual analysis (96.33%). This is the routine, time-intensive work that consumes researchers' days.

Why this is a structural shift, not a benchmark bump

Six years ago, AlphaFold could predict a structure. Today, a general model runs the whole design workflow and beats expert teams on a contested public target. Three structural consequences stand out.

First, the bottleneck moves from computation to validation. Design that once took months now takes 24-48 hours; the binding constraint is wet-lab turnaround, not thinking. Second, specialized expertise becomes optional at the front of the pipeline — orchestration is now the model's job, which is why Anthropic has started an access program for scientists even while it keeps these biological capabilities out of public-facing Claude over dual-use concerns. Third, the playbook is open: Anthropic published the prompt, data, and 1,440 designs on Hugging Face, so any lab can reproduce the workflow rather than treat the result as a black box.

For the broader pattern, see how other AI systems are pushing into discovery and screening — from virtual screening models that find real hits to AI biologists proposing drug targets.

What researchers should do now

Binder design is not yet a drug — toxicology and clinical trials still take years, and every structure here is computationally predicted, not experimentally solved. But the workflow is reproducible today, so the practical moves are clear:

  • Run the published prompt first. The 16,000-word prompt and full dataset are on Hugging Face; pilot it against a target you already have wet-lab data for, so you can measure the uplift yourself.
  • Treat orchestration as the deliverable. The durable skill is encoding expert workflows as prompts plus a tool list — that pattern transfers to any lab process, not just proteins.
  • Keep validation independent. The result only meant something because two outside labs tested it. Build a screening partner into the loop before you trust any ranking.
  • Watch the policy track. Anthropic's scientist access program and its decision to gate biological capabilities will shape whether this stays a lab advantage or becomes a general research utility.

Leave a Comment

Scroll to top