While make-on-demand libraries now span trillions of molecules, full library docking struggles beyond a few billion, motivating prioritization that recovers top-scoring compounds while evaluating only a fraction of a library. Here we introduce a similarity-based prioritization approach, ChemSTEP, and define the effective size of a library treated by any prioritization algorithm, Neff. ChemSTEP docks a representative seed set, selects diverse high-scoring “beacons”, and iteratively traverses the library through cycles of beacon selection, similarity search, and docking. Retrospectively on eight targets, ChemSTEP recovered over 75% of high-scoring compounds while docking less than 5% of a library. We then tested ChemSTEP prospectively against AmpC β-lactamase using a 13.2 billion molecule library. Because AmpC recognizes negatively charged inhibitors, we explicitly docked all 360 million library anions, synthesizing and testing 241 high-ranking ones in parallel to the ChemSTEP 13.2B run. Compared with previous docking of 99 million and 1.7 billion molecules against AmpC, the 13.2 billion library had higher hit-rates (2% vs 25% vs 37%, respectively) and found more potent compounds. Meanwhile, ChemSTEP retrieved 80% of the 241 high-ranking compounds within the first 0.5% docked. Trillion-molecule libraries might be in reach with this approach.
Olivier Mailhot, Katie L. Holland, Lu Paris et al.· ACS Central Science· 0 citations
A quantitative framework for understanding how docking performance responds to methodological improvements has been lacking is developed by modeling large-scale experiments from three previously published docking campaigns, providing an objective basis for benchmarking and comparing virtual screening approaches.
Laust Moesgaard, Brian K. Shoichet, Olivier Mailhot· Journal of Chemical Informat...· 1 citation