Skip to content

Flowr.root – A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction

Oct 2025 · arXiv.org · Vol 17 · 5 citations · ⚡ 1 influential · 125 references
Biology Computer Science Medicine

TL;DR

Flowr.root achieves state-of-the-art performance in both unconditional 3D molecule and pocket-conditional ligand generation, producing geometrically realistic, low-strain structures with computational efficiency on established benchmark datasets.

Abstract

We present Flowr.root, an SE(3)-equivariant flow-matching model for pocket-aware 3D ligand generation with joint binding affinity prediction and confidence estimation. The model supports multiple design modes including de novo generation, interaction/pharmacophore-conditional sampling, fragment elaboration, and multi-endpoint affinity prediction (pIC50, pKi, pKd, pEC50). Training combines large-scale ligand libraries with mixed-fidelity protein–ligand complexes, followed by refinement on curated co-crystal datasets and adaptation to project-specific data through parameter-efficient finetuning. Flowr.root achieves state-of-the-art performance in both unconditional 3D molecule and pocket-conditional ligand generation, producing geometrically realistic, low-strain structures with computational efficiency on established benchmark datasets. The integrated affinity prediction module demonstrates superior accuracy on the Spindr test set and outperforms recent models on the Schrödinger FEP+/OpenFE benchmark while offering substantial speed advantages. As a foundation model, Flowr.root requires continuous parameter-efficient finetuning on project-specific datasets to account for unseen structure-activity landscapes, which we demonstrate yields strong correlation with experimental in-house data. The model’s joint generation and affinity prediction capabilities enable inference-time scaling through importance sampling, effectively steering molecular design toward higher-affinity compounds. Case studies validate this approach: selective CK2α ligand generation against CLK3 shows significant correlation between predicted and quantum-mechanical binding energies, while scaffold elaboration studies on ERα, TYK2 and BACE1 demonstrate strong agreement between predicted affinities and QM calculations. By integrating structure-aware generation, affinity estimation, and property-guided sampling within a unified framework, Flowr.root provides a comprehensive foundation for structure-based drug design spanning hit identification through lead optimization.

View source

Similar papers

Open access Jul 2026

FLOWR.ROOT – A flow matching-based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction

FLOWR.root is presented, an SE(3)-equivariant flow-matching framework that jointly generates high-quality 3D ligands and predicts multi-endpoint binding affinities, demonstrating state-of-the-art performance, efficient domain adaptation, and practical impact across de novo design, scaffold elaboration, and lead optimization workflows.

Julian Cremer, Tuan Le, M. Ghahremanpour et al. · 1 citation
Preprint Jul 2026

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

Thomas MacDougall, Maksim Kuznetsov, Roman Schutski et al. · 1 citation
Open access Aug 2026

Enhanced Line Search Improves Robustness and Efficiency of Pose Sampling in Protein–Ligand Docking

Physics-based protein–ligand docking critically depends on efficient pose sampling, yet established sampling and local refinement algorithms can be inefficient and unstable in the highly nonconvex energy landscapes characteristic of protein–ligand interactions. To address this limitation, we introduce an enhanced local optimization strategy based on curved line search (CLS) and integrate it into AutoDock Vina, resulting in Vina_CLS. The proposed method enables more flexible step-size selection during local refinement and improves convergence in challenging regions of the energy landscape. Across benchmarks on the PDBbind refined set and the LEADS-PEP data set, Vina_CLS consistently outperforms the baseline, exhibiting greater robustness by solving more docking problems, as well as improved efficiency through reduced function and gradient evaluations and shorter runtimes. These gains translate into practical benefits, including more frequent identification of difficult-to-access local minima, enhanced redocking accuracy, and increased recovery of near-native poses. Together, these results demonstrate that improved local optimization can substantially enhance docking performance, highlighting an important, underexplored opportunity to advance structure-based drug discovery.

Leo Gaskin, Matthias Welsch, J. Kirchmair et al. · 0 citations
Preprint Jul 2026

Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization

Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design. Many such models follow the structure-based drug design (SBDD) paradigm, generating molecules to fit a target binding pocket. However, existing diffusion-based SBDD methods typically couple pocket and ligand representation learning, model interactions only at the atom level, and prioritize binding affinity over other developability properties. Here, we introduce conDitar-dev, a conditional diffusion-based SBDD framework for generating ligands with strong binding affinities and favorable ADMET properties. It consists of three modules: msPRL, a pretrained multi-scale pocket representation learning module; conDitar, a pocket-conditioned diffusion model guided by msPRL representations; and paOPT, a generation-time method for optimizing ligand developability. On a newly curated benchmark of human disease targets, conDitar outperforms state-of-the-art SBDD baselines, achieving an average binding score of -8.85 kcal/mol. Across five ADMET properties, conDitar-dev improves performance by up to 73% over conDitar. To further validate the abilities of conDitar-dev to generate developable molecules, we have applied it to two validated druggable targets: programmed death-ligand 1 (PD-L1) and colony-stimulating factor 1 receptor (CSF1R) proteins. Top-ranked generatively designed molecules and their analogs have been experimentally synthesized and biologically tested. Two molecules generated directly by conDitar-dev for PD-L1 exhibited SPR-derived $K_D$ values of 3.49 and 3.75 $\mu$M, respectively. Hit expansion based on conDitar-dev-designed molecules identified selective CSF1R inhibitors with IC$_{50}$ values as low as 200 nM, while also uncovering opportunities for drug repositioning.

Ruoxi Gao, Jiangweizhi Peng, Ziqi Chen et al. · 0 citations
Preprint Aug 2026

Packora: Systematic Design for Generative Molecular Crystal Structure Prediction

Packora is presented, a flow-based generative model for molecular CSP that jointly predicts atomic coordinates and the lattice from molecular graphs that outperforms the baselines on both structure generation and ranking benchmarks.

Nayoung Kim, Kiyoung Seong, Sungsoo Ahn · 0 citations
Open access Aug 2026

DASH: A Pocket-Aware and Objective-Aware Framework for Million-Scale Structure-Based Molecular Generation

Structure-based molecular diffusion models have shown considerable potential for de novo drug design. However, their practical use in million-scale candidate-library construction remains limited by fixed target-agnostic inference settings, insufficient support for objective-aware molecular prioritization, and limited integration of scalable property evaluation with standardized library production. Here, we present DASH, a pocket-aware and objective-aware framework for converting protein-conditioned diffusion outputs into property-tagged molecular libraries for computational prioritization. DASH combines pocket-complexity-aware sampling, configurable objective-aware molecular scoring, and a scalable production layer. The sampling strategy adjusts inference effort according to geometric and physicochemical features of the target binding pocket. The Objective-Aware Quality Module (OQM) filters and reranks generated molecules using configurable descriptors, desirability functions, and scoring profiles. The production layer supports million-scale execution through asynchronous GPU–CPU processing, streaming output, molecular scoring, and SDF annotation. We evaluated DASH through pocket-complexity analysis, OQM profile analysis, execution benchmarks, and multi-GPU/multinode scaling experiments. The results show that DASH can adapt inference effort across pockets, rerank generated libraries under different design objectives, and produce standardized molecular libraries for downstream analysis. An EGFR-oriented case study further demonstrates how DASH-generated libraries can support downstream computational prioritization and representative candidate selection. Together, these results demonstrate that DASH extends protein-conditioned diffusion models from raw molecular generation toward practical-scale, objective-aware candidate-library construction and computational hit prioritization.

Baohua Zhang, Huangchao Xu, Xiaoning Wang et al. · 0 citations