Aug 2026· Journal of Chemical Information and Modeling· 0 citations· 88 references
TL;DR
PockLigGPT achieves competitive docking-oriented performance under a standardized evaluation protocol while maintaining chemical plausibility, favorable physicochemical profiles, and Lipinski-based drug-likeness.
Abstract
De novo drug design aims to generate molecules targeting specific protein pockets while retaining chemical plausibility and drug-like properties. Recent 3D structure-based generative methods explicitly model pocket-ligand geometry, but this does not always translate into chemically realistic or practically usable candidate molecules. Molecular language models provide a complementary sequence-based alternative. However, it remains unclear whether sequence-based pocket information can effectively guide ligand generation, whether multistage training improves pocket-specific generation, and whether docking-guided reinforcement learning can be integrated into a practical generation pipeline. We introduce PockLigGPT, a GPT-based framework for pocket-sequence-conditioned molecular generation. Rather than producing fixed 3D coordinates, PockLigGPT formulates ligand design as a sequence-generation problem conditioned on the amino acid composition of the protein pocket. The model is trained in four stages: large-scale chemical pretraining based on ZINC20; bioactivity-oriented adaptation based on ChEMBL; pocket-sequence-conditioned fine-tuning using binding-pocket amino acid sequences paired with ligands; and, finally, pocket-specific docking-guided reinforcement learning using AutoDock Vina-based rewards. PockLigGPT achieves competitive docking-oriented performance under a standardized evaluation protocol while maintaining chemical plausibility, favorable physicochemical profiles, and Lipinski-based drug-likeness. Docking studies on Alzheimer’s disease-associated targets and token-level analyses further support the utility of PockLigGPT for de novo drug design.
Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design. Many such models follow the structure-based drug design (SBDD) paradigm, generating molecules to fit a target binding pocket. However, existing diffusion-based SBDD methods typically couple pocket and ligand representation learning, model interactions only at the atom level, and prioritize binding affinity over other developability properties. Here, we introduce conDitar-dev, a conditional diffusion-based SBDD framework for generating ligands with strong binding affinities and favorable ADMET properties. It consists of three modules: msPRL, a pretrained multi-scale pocket representation learning module; conDitar, a pocket-conditioned diffusion model guided by msPRL representations; and paOPT, a generation-time method for optimizing ligand developability. On a newly curated benchmark of human disease targets, conDitar outperforms state-of-the-art SBDD baselines, achieving an average binding score of -8.85 kcal/mol. Across five ADMET properties, conDitar-dev improves performance by up to 73% over conDitar. To further validate the abilities of conDitar-dev to generate developable molecules, we have applied it to two validated druggable targets: programmed death-ligand 1 (PD-L1) and colony-stimulating factor 1 receptor (CSF1R) proteins. Top-ranked generatively designed molecules and their analogs have been experimentally synthesized and biologically tested. Two molecules generated directly by conDitar-dev for PD-L1 exhibited SPR-derived $K_D$ values of 3.49 and 3.75 $\mu$M, respectively. Hit expansion based on conDitar-dev-designed molecules identified selective CSF1R inhibitors with IC$_{50}$ values as low as 200 nM, while also uncovering opportunities for drug repositioning.
Ruoxi Gao, Jiangweizhi Peng, Ziqi Chen et al.· 0 citations
A clear pattern is revealed in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.
Thomas MacDougall, Maksim Kuznetsov, Roman Schutski et al.· 1 citation
This work proposes LLMol, a principled reinforcement learning framework that directly incorporates verifiable rewards for targeted molecule generation and introduces Reinforcement Learning with Verifiable Rewards (RLVR), which directly integrates property-based reward signals to guide molecular generation toward task-specific objectives.
MolecularCanvas is an interactive system that enables users to iteratively construct an optimization context by integrating high-level goals, structure-level annotations, property constraints, and reference-based preferences that guides the generation of candidate molecules across diverse molecular structures.
Haoyu Dong, Rui Sheng, Shuhao Zhang et al.· 0 citations
PGFS++ is introduced, a synthesis-aware reinforcement learning framework for input-specific molecular improvement that improves target properties while preserving high output diversity, and experiments show that PGFS++ improves target properties while preserving high output diversity.
Boqiao Zhang, Godbless James, S. Gottipati et al.· 0 citations
The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning and outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016.