BoltzMol-1: Virtual Screening That Finds Real Hits

Most virtual screening papers sell you a score: ROC-AUC, enrichment factor, docking rank. BoltzMol-1 sells something harder to fake — wet-lab receipts. The team behind Boltz-2, the protein–small-molecule structure model, ran a prospective screen across 10 diverse targets, buying and testing just 28–96 compounds per target, and came back with binders or functional actives on 6 of them. That flips the metric of success from “how high did the model rank” to “how many of the compounds we actually bought worked.”

From benchmark scores to prospective hit rates

Retrospective benchmarks are necessary but not sufficient. A model that scores well on a held-out split of known binders can still fail in a real project, where the target lacks a co-crystal structure, the pocket is a shallow surface or an allosteric site, the commercial catalog is full of poor starting points, and the budget cannot cover thousands of compounds. Traditional HTS finds real actives, but hit rates often sit at 0.01–0.14% and the cost is brutal. Classic docking scales, yet its scoring functions wobble on GPCR allosteric pockets, protein–protein interfaces and voltage-sensing domains. BoltzMol-1’s answer is not a bigger screen; it is a tighter workflow: shrink the candidate space with the model, filter out molecules that are not worth testing with medicinal chemistry and ADMET rules, then buy only a few dozen compounds and put them in real assays.

Seven steps from sequence to purchase list

  • Define the target — protein sequence, chain ID, optional pocket residues or a reference ligand.
  • Prepare the candidate library from purchasable or synthesizable spaces (WuXi off-the-shelf, Enamine REAL).
  • Medicinal chemistry pre-filter — PAINS, reactive groups, extreme property profiles.
  • Virtual PAINS exclusion — remove molecules the model loves across unrelated targets.
  • Co-fold and score with Boltz-2 — structure, affinity and binding-confidence signals.
  • Composite score ranking — combine multiple Boltz-2 outputs into one rank.
  • ADMET triage and purchase — apply logS, logD and Caco-2 predictions, then buy the top N within budget.

For fixed enumerable libraries the pipeline scores and ranks everything; for combinatorial spaces like Enamine REAL it switches to generative active learning, iterating proposals to concentrate the search in promising regions.

Virtual PAINS: a counter-screen for the model itself

Real PAINS compounds are assay artifacts — aggregators, reactive or fluorescent molecules that fake activity. BoltzMol-1 introduces the model-side equivalent: virtual PAINS. Some molecules get high scores from Boltz-2 across many unrelated targets, not because they bind everything but because they exploit a scoring bias. The team aggregates scores over a set of held-out targets, flags molecules that consistently rank near the top regardless of target, and removes them and their close analogs from the pool. It is counter-screen thinking moved into the ranking stage: if a compound looks great against every target, it may actually be a probe of the model’s blind spots.

ADMET before you buy

With a budget of a few dozen compounds per target, every slot is precious. BoltzMol-1 brings developability into hit selection instead of bolting it on later. Three internal models: logS (binned into high/low for confident triage rather than a fragile continuous value), logD, and Caco-2 A→B permeability, trained on ChEMBL and GOSTAR with leakage controls that check matched molecular pairs, not just whole-molecule similarity. The networks are GPS-style graph transformers mixing local message passing with global atom-pair attention and per-atom pKa features. On prospective profiling, predicted logD hit r = 0.91 against experiments; the low-solubility bin correctly flagged 29 of 34 compounds with zero false “high solubility” calls; Caco-2 reached r = 0.65 with efflux inhibitors in the assay, and the paper admits the model cannot reliably identify efflux substrates. Triage-grade, not DMPK-grade — which is exactly the right framing for hit selection.

The lab results: 6 of 10 targets

  • GLP-2R (class B GPCR with no small-molecule co-structure): 12 functional antagonists from 51 compounds, IC50 0.07–24 µM, with no agonist/antagonist bias baked into the screen.
  • MRGPRX2 (shallow, solvent-exposed GPCR): 10 agonists plus 3 antagonists from 38 compounds.
  • ROR1 (pseudokinase): a confirmed binder across SPR, MST and nanoDSF, SPR KD 5.13 µM, from 34 compounds.
  • LC3B/GABARAP (PPI surface): 3 peptide-displacement inhibitors at IC50 15–40 µM, from 28 compounds sampled from Enamine REAL.
  • PknB (TB kinase): 16 actives from 96 compounds, IC50 23.2–399 µM — high hit rate, modest potency.
  • STAT6 (SH2 domain, PPI-like): 1 confirmed binder at KD 30.3 µM, counter-screened against p38.

The headline is not that the model solved virtual screening. It is that on a low budget, a model-plus-filters-plus-procurement workflow can generate real, testable starting points across hard target classes — and that functional mode (agonist vs antagonist) is still an experimental readout, not a prediction.

What the four failures teach

mGlu4 PAM, AMY3, MALT1 and Nav1.8 came back empty — and that is information. Nav1.8 is the cleanest lesson: the target family has strong sequence coverage in training data, but the screen aimed at the VSD2 allosteric site while known small-molecule structures mostly sit in the pore domain. Sequence similarity did not transfer to binding-mode similarity. mGlu4 PAM depends on subtle conformational states; AMY3 is a multi-subunit receptor complex; MALT1 failed despite rich known ligands, a reminder that in-distribution experience does not guarantee a new chemotype, and that commercial catalog bias can starve a target of the right chemistry. The takeaway: target setup — conformation, site definition, complex assembly — can matter more than the model itself.

Why this is an industry signal

The paper is a marker, not a finish line. Evaluation in AI drug discovery is shifting from benchmark leaderboards to prospective hit rate: how many compounds tested per target, what fraction hit, whether hits are purchasable and optimizable, and what a confirmed hit costs. That matches a broader pattern — AI research tools are becoming workflows rather than single scorers, from open-source research agents that run literature-to-paper pipelines to AI systems that organize themselves. The practical consequence: teams without massive HTS budgets get a lower-cost entry into hit discovery, and even micromolar binders can serve structural biology as crystallization or cryo-EM starting points. None of this removes the medicinal chemist or the structural biologist; it moves their attention earlier in the pipeline.

What to do next

  • When evaluating a screening model, ask for prospective data, not just AUC: per-target budgets, hit rates, orthogonal assays, counter-screens.
  • Put ADMET triage before procurement, not after binding.
  • Treat functional activity as an experiment: pair screens with state-selective or functional readouts when the target class demands it.
  • Add your own virtual PAINS step — cross-target score aggregation is cheap and filters model bias.
  • Track cost per confirmed hit and time from sequence to lab result; those are the numbers that actually improve.

Also read the limitations: composite score weights are not fully disclosed, there is no head-to-head against a strong docking workflow under identical budgets, and most hits are early-stage micromolar binders rather than leads. That transparency is itself a sign the field is maturing.

Leave a Comment

Scroll to top