Skip to content

Category

natural language processing

2,394 papers

#natural language process... Open access Nov 2025

A virtual platform for automated hybrid organic-enzymatic synthesis planning

The results indicate that this fully automated, open-source system holds potential value for improving the efficiency and sustainability of molecular synthesis, and the integration of organic and enzymatic synthesis enhances molecule construction efficiency.

Xiaorui Wang, Xiaodan Yin, Xujun Zhang et al. · 0 citations
#machine learning Open access Nov 2025

mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset

The authors develop mRNABERT, a foundational AI model that designs entire mRNA sequences and demonstrates superior performance across comprehensive benchmarks, which signifies a substantial leap forward in mRNA research and therapeutic development.

Ying Xiong, Aowen Wang, Yu Kang et al. · 23 citations · ⚡1
#machine learning Open access May 2025

Token-Mol 1.0: tokenized drug design with large language models

Token-Mol is presented, a token-only 3D drug design model that encodes both 2D and 3D structural information, along with molecular properties, into discrete tokens, which introduces a Gaussian cross-entropy loss function tailored for regression tasks, enabling superior performance across multiple downstream applications.

Ji-Ke Wang, Rui Qin, Mingyang Wang et al. · 30 citations · ⚡1
#natural language process... Open access Oct 2025

Committors without Descriptors

The study of rare events is one of the major challenges in atomistic simulations, and several enhanced sampling methods toward its solution have been proposed. Recently, it has been suggested that the use of the committor, which provides a precise formal description of rare events, could be of use in this context. We have recently followed up on this suggestion and proposed a committor-based method that promotes frequent transitions between the metastable states of the system and allows extensive sampling of the process transition state ensemble. One of the strengths of our approach is being self-consistent and semiautomatic, exploiting a variational criterion to iteratively optimize a neural-network-based parametrization of the committor, which uses a set of physical descriptors as input. Here, we further automate this procedure by combining our previous method with the expressive power of graph neural networks, which can directly process atomic coordinates rather than descriptors. Besides applications on benchmark systems, we highlight the advantages of a graph-based approach in describing the role of solvent molecules in systems, such as ion pair dissociation or ligand binding.

Peilin Kang, Jintu Zhang, Enrico Trizio et al. · 6 citations
#machine learning Open access Jul 2025

RSGPT: a generative transformer model for retrosynthesis planning pre-trained on ten billion datapoints

RSGPT, a generative model pre-trained on ten billion data points, achieving state-of-the-art performance for synthesis planning, and introduces reinforcement learning to capture the relationships among products, reactants, and templates more accurately.

Yafeng Deng, Xinda Zhao, Hanyu Sun et al. · 18 citations · ⚡2
#machine learning Open access Apr 2026

Accurate and task-agnostic modeling of enzymatic reactions through multimodal relational learning

ERAM aligns pre-trained molecular representations from Protein Language Model with the knowledge of enzyme catalysis by modeling enzymatic reactions as multi-relational data, and demonstrates its potential as a versatile and effective tool for enzyme catalysis research.

Yuansheng Huang, Lanqing Li, Wenjia Qian et al. · 2 citations
#natural language process... Open access Apr 2026

LaMGen: LLM-based 3D molecular generation for multi-target drug design

Multi-target drugs hold great promise for treating complex diseases, yet existing methodologies predominantly rely on ligand-based approaches, which lack sufficient biological context and are often confined to specific target pairs, resulting in limited generalizability. Here, we introduce LaMGen, a general-purpose multi-target drug design framework powered by large language models (LLMs). Built on MTD2025, a dataset comprising over 600,000 quantum-accurate molecular conformations and 700,000 multi-target associations, LaMGen directly yields energy-favorable conformations with quantum-level accuracy. The framework integrates ESM-C protein embeddings, rotation-aware ligand tokens, and a TriCoupleAttention module to capture multi-level target–ligand interactions. Across independent benchmarks, LaMGen outperforms diffusion-based model across multiple properties, generating molecules in an average of 0.44 s, while preserving high conformational plausibility. Retrospective analyses demonstrate that LaMGen not only can reproduce molecules identical to known actives, but also consistently produces structurally novel candidates with conserved core scaffolds and superior binding affinities. Designing effective multi-target therapeutics remains a major challenge, as existing ligand- or protein-centric methods struggle to generate biologically contextualized, spatially valid 3D molecules, particularly for triple-target systems. This study introduces LaMGen, an LLM-powered framework that leverages large-scale protein-ligand data and rotation-aware molecular encoding to rapidly produce chemically plausible multi-target candidates, achieving strong zero-shot generalization, superior molecular quality, and robust performance across dual- and triple-target design tasks.

Qun Su, Qiaolin Gou, Hui Zhang et al. · 1 citation
#natural language process... Preprint Aug 2026

DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation

Prompting-based (i.e., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not changed for the task, face three problems: (i) relying on coarse-grained schema information that may not reveal the fine-grained relationships needed to distinguish ambiguous columns, (ii) failing to capture recurring SQL-generation failures, and (iii) suffering from omission or hallucination of components in complex questions. This paper develops DexterSQL, a prompting/non-fine-tuning-based Text-to-SQL system that improves SQL generation with three novel components: (i) deep schema explorator that identifies ambiguous columns, analyzes their individual and joint data distributions to uncover their relationships and the distinct role of each, (ii) database-agnostic rule creator that mines mismatches between generated and gold SQL only on the training database and converts them into database-agnostic corrective rules that capture recurring LLM failure patterns; and (iii) multi-path SQL generation that introduces a dependency-tree-based intermediate representation that uses the question's sentence structure to guide its decomposition into an SQL skeleton for final SQL generation. DexterSQL achieves a higher accuracy compared to the state-of-the-art using both open-source/weight and closed-source/weight models. Particularly, DexterSQL shows a high improvement of at least 5.5% using an open-weight model (GPT-OSS-120B) on BIRDDev, with total accuracy 70.4%. DexterSQL also shows better improvement of at least 1.4% using closed-weight models, with total accuracy 72.1% and 72.9% on BIRD-Dev with GPT-4o and GPT-5.2.

Anik Pramanik, Murat Kantarcioglu, Vincent Oria et al. · 0 citations
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#computer vision Preprint Aug 2026

AI Sandbox: Technical Report

This work presents the design and implementation of a governance-aware, multi-tenant AI sandbox for structured experimentation and the generation of reusable evaluation evidence across projects and stakeholder groups.

Muhammad Waseem, M. Islam, Md Nasir Uddin Shuvo et al. · 0 citations
#computer vision Review Feb 2026

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review

A Multi-Vocal Literature Review is conducted, combining insights from both academia and industry, including peer-reviewed studies and grey literature to systematically synthesize and analyze existing knowledge on LLM-based multi-agent systems for code generation.

Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al. · 2 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

MIT News · Artificial Intelligence Aug 20, 2026

Paving the way for greener ammonia production

New MIT research could lead to better materials for a fossil-fuel-free process for making the chemical that's essential to fertilizer and other products.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.