Overview
In-stock screening collections top out around ten million compounds. Make-on-demand combinatorial spaces — Enamine REAL and comparable ultra-large libraries — contain tens of billions, and they grow faster than any docking fleet can screen them. Brute force loses here: even at GPU speeds, docking a 36-billion-compound space is not a bigger version of the same job, it is a different problem.
The way through is chemical, not computational. An ultra-large library is not a list of molecules; it is a set of reactions and their building blocks (synthons). Every product is reachable by plugging building blocks into a reaction pattern. That structure is exploitable: instead of docking finished products, you dock each building block once as a small capped fragment, identify which fragments bind productively with room to grow into the pocket, and then use reaction enumeration to generate only the products of those winners. This is synthon screening — the approach introduced by V-SYNTHES2 [1] — and it converts an infeasible screening campaign into a tractable one, with the published method reporting compute reductions on the order of 10,000×.
Autochem implements the full loop — fragment library generation, fragment docking, geometric selection, targeted enumeration, product docking, and post-processing — as agent-orchestrated pipelines on the same infrastructure as our docking protocol. Every molecule we generate carries complete lineage: which reaction produced it, from which building blocks, descending from which docked fragment. Nothing in a billion-compound space is ever an anonymous row.
Pipeline Steps
Step 1: Chemical Space Ingestion
We ingest reaction definitions (reaction SMARTS plus per-position building-block lists) from three sources: vendor ultra-large libraries under the customer's license, the customer's proprietary reaction and synthon collections, or spaces we curate and build together with you. The importer classifies each reaction automatically — two-component versus three-component, and bridge-like topologies where a central synthon must be handled in iteration order — because the enumeration strategy and the fragment library shape both depend on that classification.
Step 2: Minimal Enumeration Library (MEL) Generation
The fragment library is built by enumerating each building block with its remaining attachment points capped by minimal placeholder groups — methyl for aliphatic chemistry, phenyl for aromatic — each carrying an isotope label so the cap position is traceable through every downstream step. The result is a Minimal Enumeration Library: one small, mostly rigid fragment per building block per reaction position, typically orders of magnitude smaller than the product space it represents. Enumeration is deterministic (seeded RDKit [4]), so a fragment library is reproducible and auditable.
Step 3: Fragment Preparation and Docking
Fragments go through exactly the same deterministic preparation chain as any library compound in our docking protocol: RDKit ETKDGv3 conformer generation, protonation at physiological pH, force-field minimization, and Meeko PDBQT conversion. Docking runs on Uni-Dock [2] against a receptor prepared by the same validated chain — structure fixing and protonation, MolProbity-checked quality, P2Rank pocket prediction. Scoring is therefore directly comparable between the fragment stage and the product stage, which is what makes the funnel's ranking transferable.
Step 4: Fragment Selection by Score and Growth Capacity
Not every well-docked fragment deserves enumeration. The selection step scores each docked fragment pose on two axes:
- Docking score. How well the fragment's core binds in its own right — the single strongest predictor of how its products will score.
- Growth capacity. A geometric analysis in the spirit of the paper's CapSelect: from the capped attachment point, a chain of probe spheres grow outward through the pocket under clearance and cone-angle constraints. A fragment whose caps point into open, druggable volume can be grown into high-affinity products; a fragment whose caps bury into protein cannot, however well it scores.
The two are combined by a weighted selection score. Weights are not constants imported from a paper — they are calibrated for the scoring engine and the target, because we have measured how strongly fragment scores inherit into product scores and how that trade behaves at each scale. This calibration step is part of the protocol, not an afterthought.
Step 5: Targeted Reaction Enumeration
Selected fragments are re-run through their reactions with the placeholder caps replaced by the real building-block lists — RDKit reaction enumeration at distributed scale. Only the selected fragments' products are generated, so the enumerated library can reach into the millions while the docking budget stays anchored at the fragment stage plus the winners. Every product inherits full lineage to its parent fragment, reaction, and synthons, which is what makes later hit triage chemically interpretable: when a product becomes a hit, you know exactly which fragment and which building blocks earned it.
Step 6: Product Preparation and Docking
Enumerated products flow through the identical preparation chain and into the two-stage Uni-Dock screen (fast pass over the enumerated set, detail re-docking with explicit hydrogens for the top candidates). Work distribution is keyed per ligand batch and pocket, so a screen of any size spreads evenly across the GPU fleet instead of concentrating on hot workers — the kind of distribution detail that decides whether a giga-scale campaign takes a day or stalls for a week.
Step 7: Iteration for Multi-Component Reactions
Three-component reactions run the selection-enumeration loop a second time: the first enumeration leaves one position capped, the products are docked, selection runs again, and only then is the final position enumerated. Bridge-like topologies, where a central synthon connects two variable positions, get the corresponding special-cased iteration order during ingestion and enumeration.
Step 8: Post-Processing and Hit Triage
The docked product set is filtered for developability — physicochemical property ranges, PAINS and reactive-group alerts, novelty against known target ligands — then clustered for diversity so the final candidate list is not a family of near-duplicates. Candidates flow directly into the rest of the platform: counterscreening for selectivity against off-target pockets and MD with MM-PBSA for physics-based re-ranking before anything is proposed for synthesis.
Validation Against a Fully Enumerated Ground Truth
Synthon screening earns trust the same way every Pauling protocol does: by running the controls. Because the method's promise is that fragment-stage results predict product-stage results, we validated it on a reaction space small enough to dock every single product — perfect ground truth, no sampling — and measured the funnel against it directly.
| Control | Result |
|---|---|
| Score inheritance — does a fragment's docking score predict its products' scores? | Spearman rank correlation of 0.91 between fragment score and product scores across all docked fragments. The premise the entire funnel rests on holds, strongly. |
| Funnel recall — docking only a fraction of the space, how many true top compounds does it find? | Docking under 8% of the space recovered more than half of the true top-100 compounds — about 7× what random selection finds at the same budget, and still 1.7× under a conservative property-matched control. |
| Pose consistency — do fragment binding modes survive enumeration? | 84% of fragments retain at least one product that reproduces the fragment's binding mode (sub-ångström median deviation for the best product per fragment). |
At benchmark scale the compute saving is about 10×; at make-on-demand scale, where the space exceeds any docking budget by orders of magnitude, the same machinery is what makes the screen possible at all — random or exhaustive selection has nothing to offer when you can only dock a sliver of the space.
Summary of Tools
| Step | Tool | Purpose | Reference |
|---|---|---|---|
| Space ingestion | Reaction importer (in-house) | Reaction/synthon ingestion and classification | — |
| Fragment library generation | RDKit | Deterministic MEL enumeration with isotope-labeled caps | [4] |
| Fragment selection | Growth-capacity analysis (in-house) | Pocket-growth geometry, CapSelect-style scoring | [1] |
| Reaction enumeration | RDKit | Targeted product enumeration from selected fragments | [4] |
| Ligand preparation | RDKit ETKDGv3, OpenBabel, Meeko | Conformers, protonation, minimization, PDBQT | — |
| Docking | Uni-Dock | GPU-accelerated fragment and product screening | [2] |
| Docking scoring function | AutoDock Vina | Scoring function underlying Uni-Dock | [3] |
| Receptor preparation | PDBFixer, Reduce, PDB2PQR, P2Rank | Same validated chain as the docking protocol | — |
| Pose quality control | PoseBusters | Physics-based pose validation | [5] |
| Distributed compute | Apache Beam | Scalable pipeline execution on GPU clusters | — |
Infrastructure
All synthon pipelines run as distributed Apache Beam jobs on elastic GPU clusters, orchestrated by the same AI agent that drives the rest of the platform. Every generated molecule, docked pose, and selection score is persisted with lineage in the platform database, so a screening campaign of any size is queryable, auditable, and re-runnable. Bring a licensed vendor space, your proprietary reaction collection, or neither — we will build the space with you.
References
- Nazarova, K. et al. (2026). V-SYNTHES2: synthon-based hierarchical virtual screening of ultra-large chemical spaces. npj Drug Discovery.https://doi.org/10.1038/s44386-026-00053-6
- Yu, Y. et al. (2023). Uni-Dock: GPU-accelerated docking enables ultralarge virtual screening. Journal of Chemical Theory and Computation, 19(11), 3336–3345.https://doi.org/10.1021/acs.jctc.2c01145
- Trott, O. and Olson, A.J. (2010). AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of Computational Chemistry, 31(2), 455–461.https://doi.org/10.1002/jcc.21334
- Landrum, G. et al. RDKit: Open-source cheminformatics.https://www.rdkit.org
- Buttenschoen, M., Morris, G.M., Deane, C.M. (2024). PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chemical Science, 15, 3130–3139.https://doi.org/10.1039/d3sc04185a