Skip to content

Category

generative ai

486 papers

#generative ai Open access Aug 2026

Adaptive-CGAN: a comprehensive framework for generative modeling on google cloud platform

Machine learning has become increasingly important in medical diagnosis, yet its effectiveness depends on access to large, reliable, and high-quality datasets. During epidemics and emerging diseases, such as COVID-19, acquiring sufficient real-world medical images rapidly is challenging. To address these issues, this study presents an Adaptive Conditional Generative Adversarial Network (Adaptive-CGAN) integrated with a cloud-based medical image processing framework. The proposed approach makes three main contributions. First, Adaptive-CGAN generates high-fidelity synthetic medical images that closely resemble real samples while improving the distinction between real and fake images. Second, a scalable TensorFlow Records (TFRecords)-based pipeline is implemented on Google Cloud Platform (GCP) to support efficient storage, loading, and processing of large-scale medical datasets. Third, a real-world COVID-19 medical image dataset comprising four disease classes is compiled and used to enhance diagnostic prediction. Experimental evaluation was conducted against several baseline generative models, including AC-GAN, WGAN, Pix2Pix, BigGAN, and CWGAN, with AC-GAN serving as the primary like-for-like baseline. Adaptive-CGAN improved classification accuracy from 93.75% ± 1.10% to 99.60% ± 0.40%, increased the Inception Score from 7.20 ± 0.28 to 8.47 ± 0.18 (+ 17.65%), reduced FID from 28.41 ± 1.35 to 26.13 ± 0.82 (− 8.02%), and reduced KID from 0.0156 ± 0.0021 to 0.0143 ± 0.0015 (− 8.33%). Adaptive-CGAN also reduced training time by 15.8% and CO₂ emissions by 9.5% compared with AC-GAN. Moreover, the GCP-based Adaptive-CGAN deployment emitted only 0.0023 kg CO₂, compared with 0.0047 kg CO₂ for the local setup, representing an approximately 50% reduction due to optimized cloud execution and the TFRecord-based pipeline. These results demonstrate that Adaptive-CGAN can improve diagnostic performance, synthetic image quality, computational efficiency, and environmental sustainability in AI-assisted healthcare.

W. Saber, A. El-Baz, R. Rizk · 0 citations
#generative ai Review Open access Aug 2026

Impact of Generative Artificial Intelligence Use on the Academic Performance of Undergraduate Students in Canada

Generative artificial intelligence (AI) has entered everyday undergraduate study faster than the evidence base has kept up with it, and the evidence that does exist points in opposite directions. This paper synthesizes peer-reviewed and carefully delimited contextual research on how generative AI use relates to undergraduate academic performance, with attention to what Canadian universities can already act on. Searches of Google Scholar, Scopus, Web of Science, ERIC, and ScienceDirect covering 2022 to 2026 produced a two-tier evidence base: Tier 1 peer-reviewed empirical and synthetic studies of student learning outcomes, and Tier 2 supplementary sources admitted under stated justifications, including contextual Canadian and international surveys, one secondary-school field experiment, and one preprint mechanistic study. Experimental syntheses report medium to large short-term gains when ChatGPT is built into instruction. Survey work points the other way: frequent unstructured use tracks with procrastination, self-reported memory problems, and slightly lower grades, and unrestricted access during practice has been shown to depress later unaided performance. Purpose of use reconciles most of that disagreement, because a tool that scaffolds thinking behaves very differently from one that replaces it. Canadian peer-reviewed studies document heavy campus adoption and considerable student ambivalence about integrity and learning, yet almost none link purpose-differentiated use to measured performance. Closing that gap matters for assessment redesign, AI-literacy programming, and the credibility of the credentials Canadian universities issue.

Pragalvha Sharma · 0 citations
#artificial intelligence Review Nov 2025

Accelerating Covalent Drug Discovery: Recent Advances in Covalent Docking Tools

Covalent inhibitors have garnered renewed attention in recent years, with their rational design becoming increasingly critical in drug discovery. Among the technologies facilitating the discovery of covalent inhibitors, covalent docking has emerged as a pivotal tool in various stages of drug development including virtual screening, lead optimization, and mechanistic studies. Since its inception as an extension of conventional docking methods in the early 2000s, covalent docking tools have undergone substantial advancements. This review provides a comprehensive overview of covalent docking algorithms, systematically categorizing their approaches according to covalent bond formation, which primarily include tethered docking, biased docking, and dynamic covalent docking approaches. A comparative analysis of current covalent docking tools is provided, alongside a critical discussion of remaining challenges. Special emphasis is placed on the growing impact of artificial intelligence (AI) in shaping novel methodologies and expanding the capabilities of covalent docking. Finally, we discuss prospects for advancing covalent docking methodologies and their applications in drug discovery.

Shi Li, Hongyan Du, Hui Zhang et al. · 2 citations
#computer vision Jan 2026

NavDB: A Comprehensive Database for Voltage-Gated Sodium Channels Modulators and Targets

Voltage-gated sodium channels (VGSCs/Navs) are essential targets for the treatment of numerous neurological, muscular, and cardiac disorders. Despite the increasing clinical interest in subtype-selective modulators, current public databases provide fragmented and inconsistent information on VGSC-related compounds and targets, particularly lacking coverage on peptides. To address this limitation, we developed NavDB, a specialized and open-access database focusing on VGSC modulators and targets. NavDB integrates 8023 curated data records covering 5168 compounds, including small molecules, toxins, drugs, and peptides, along with comprehensive annotations on biological activity, druggability, and structural feature. NavDB also features advanced functions such as text-based and structure-based search, peptide similarity matching, and AI-powered property prediction. Moreover, the database offers high-quality 3D visualizations of targets and peptides, with disulfide bond and signal peptide annotations. All data are freely downloadable to support both experimental and computational drug discovery. NavDB is publicly available at: http://cadd.zju.edu.cn/navdb/.

Gaoang Wang, Jiahui Yu, Haiyi Chen et al. · 0 citations

Computational and AI-Driven Ecosystem for Structure-Based Covalent Drug Discovery.

ConspectusThe field of covalent drug discovery has witnessed a remarkable resurgence in recent years, a trend underscored by the approval of more than 125 covalent drugs by the US FDA as of 2025, which demonstrates their immense therapeutic potential. Driven by ever-increasing computational power and vast amounts of data, deep learning (DL) is profoundly transforming numerous fields, from natural language processing to drug discovery. In the development of covalent drugs, in particular, advanced computational methods centered on data-driven approaches and artificial intelligence (AI) exhibit immense potential. The realization of this potential depends on the construction of a synergistic ecosystem. Here, we define this "ecosystem" as an integrated set of components─including (i) curated covalent-relevant databases, (ii) AI/physics-based predictive and scoring models, (iii) interoperable computational workflows spanning site identification, docking/virtual screening, and lead optimization, and (iv) closed-loop feedback that systematically incorporates experimental outcomes to update data resources and refine/validate models. This begins with the systematic collection of past experimental results to build high-quality databases. These databases, in turn, provide the foundation for developing AI-driven computational tools capable of precisely interfacing with and accelerating downstream tasks, such as molecular docking (for generating physically plausible conformations and conducting large-scale virtual screening) and lead optimization. The application of these AI tools not only guides experimental design, but the resulting key data also feed back into and enrich the databases. Furthermore, in the cutting-edge field of covalent drugs, the precise identification of "druggable" covalent sites on target proteins has emerged as another critically important downstream task.In this Account, we describe a computational and AI-driven ecosystem for structure-based covalent drug discovery and highlight our contributions to this field. By explicitly linking databases, models, workflows, and experimental feedback into a single framework, this Account moves beyond a simple inventory of individual tools to instead offer a systematic and panoramic perspective on an integrated ecosystem for covalent drug discovery, driven by data and computational engines including AI. We focus on how this ecosystem systematically addresses the challenges from covalent binding site identification to lead discovery, thereby fundamentally accelerating the development of next-generation covalent therapies. We first articulate the philosophy behind the construction and updating of covalent databases, emphasizing the necessity of high-quality data. Subsequently, we delve into a suite of cutting-edge, AI-driven computational methods, exploring the potential of deep learning in tasks such as molecular docking, covalent binding site prediction, and lead optimization. To bridge the gap between computational theory and experimental validation, we will use the discovery of potent covalent CRM1 inhibitors as a specific case study, detailing how our customized, structure-based virtual screening pipeline was utilized to achieve a seamless workflow from computational prediction to biological validation. This section is intended to offer actionable guidance for experimental researchers seeking to leverage these powerful computational tools. Finally, we highlight the limitations and potential pitfalls of this AI engine─concerns that are equally relevant when developing AI-driven covalent docking algorithms. Building on our group's recent benchmarking of AI docking methods, we objectively evaluate current performance and discuss how transformative advances such as AlphaFold3 may reshape the field.

Shi Li, Hongyan Du, Xujun Zhang et al. · 4 citations
#natural language process... Open access Nov 2025

A virtual platform for automated hybrid organic-enzymatic synthesis planning

The integration of organic synthesis with enzymatic catalysis offers a promising route toward efficient and sustainable construction of complex molecules. While organic synthesis enables diverse transformations, enzymatic catalysis enhances stereoselectivity under mild conditions, improving cost-effectiveness and environmental impact. However, current enzymatic synthesis planning algorithms face challenges in formulating robust hybrid organic–enzymatic strategies. Key issues include the difficulty in devising hybrid planning approaches and the reliance on template-based enzyme recommendations, which limits their adaptability across diverse scenarios. Here we show ChemEnzyRetroPlanner, an open-source hybrid synthesis planning platform that combines organic and enzymatic strategies with AI-driven decision-making. The platform features advanced computational modules, including hybrid retrosynthesis planning, reaction condition prediction, plausibility evaluation, enzymatic reaction identification, enzyme recommendation, and in silico validation of enzyme active sites. A central innovation is the RetroRollout* search algorithm, which outperforms existing tools in planning synthesis routes for organic compounds and natural products across multiple datasets. ChemEnzyRetroPlanner provides an intuitive graphical interface and programmatic APIs for scalability, while leveraging the chain-of-thought strategy and the Llama3.1 model to autonomously activate hybrid synthesis strategies for diverse scenarios. The results indicate that this fully automated, open-source system holds potential value for improving the efficiency and sustainability of molecular synthesis. The integration of organic and enzymatic synthesis enhances molecule construction efficiency. Here, the authors present ChemEnzyRetroPlanner, an AI-driven platform for automated hybrid synthesis planning, improving synthesis route efficiency and sustainability.

Xiaorui Wang, Xiaodan Yin, Xujun Zhang et al. · 0 citations
#machine learning Open access Nov 2025

mRNABERT: advancing mRNA sequence design with a universal language model and comprehensive dataset

Designing effective mRNA sequences for therapeutics remains a formidable challenge. Inspired by successes in protein design, language models (LMs) are now being applied to RNA, but progress is often impeded by the lack of comprehensive training data. Existing models are frequently limited to UTR or CDS regions, restricting their application for complete mRNA sequences. We introduce mRNABERT, a robust, all-in-one mRNA designer pre-trained on the largest available mRNA dataset. To enhance performance, we propose a dual tokenization scheme with a cross-modality contrastive learning framework to integrate semantic information from protein sequences. On a comprehensive benchmark, mRNABERT demonstrates state-of-the-art performance, outperforming previous models in the majority of tasks for 5’ UTR and CDS design, RNA-binding protein (RBP) site prediction, and full-length mRNA property prediction. It also surpasses large protein models in several related tasks. In conclusion, mRNABERT’s superior performance across these diverse tasks signifies a substantial leap forward in mRNA research and therapeutic development. Designing complete mRNA sequences for new vaccines and therapies is a complex challenge. Here, the authors develop mRNABERT, a foundational AI model that designs entire mRNA sequences and demonstrates superior performance across comprehensive benchmarks.

Ying Xiong, Aowen Wang, Yu Kang et al. · 22 citations · ⚡1
#machine learning Open access Sep 2025

Unified and explainable molecular representation learning for imperfectly annotated data from the hypergraph view

Molecular representation learning (MRL) has shown promise in accelerating drug development by predicting chemical properties. However, imperfectly annotation among datasets pose challenges in model design and explainability. In this work, we formulate molecules and corresponding properties as a hypergraph, extracting three key relationships: among properties, molecule-to-property, and among molecules, and developed a unified and explainable multi-task MRL framework, OmniMol. It integrates a task-related meta-information encoder and a task-routed mixture of experts (t-MoE) backbone to capture correlations among properties and produce task-adaptive outputs. To capture underlying physical principles among molecules, we implement an innovative SE(3)-encoder for physical symmetry, applying equilibrium conformation supervision, recursive geometry updates, and scale-invariant message passing to facilitate learning-based conformational relaxation. OmniMol achieves state-of-the-art performance in properties prediction, reaches top performance in chirality-aware tasks, demonstrates explainability for all three relations, and shows effective performance in practical applications. Our code is available in our https://github.com/bowenwang77/OmniMol public repository. AI models for drug discovery often struggle with real-world, incomplete data. Here, the authors present OmniMol, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability.

Bowen Wang, Junyou Li, Donghao Zhou et al. · 9 citations
#machine learning Open access May 2025

Token-Mol 1.0: tokenized drug design with large language models

The integration of large language models (LLMs) into drug design is gaining momentum; however, existing approaches often struggle to effectively incorporate three-dimensional molecular structures. Here, we present Token-Mol, a token-only 3D drug design model that encodes both 2D and 3D structural information, along with molecular properties, into discrete tokens. Built on a transformer decoder and trained with causal masking, Token-Mol introduces a Gaussian cross-entropy loss function tailored for regression tasks, enabling superior performance across multiple downstream applications. The model surpasses existing methods, improving molecular conformation generation by over 10% and 20% across two datasets, while outperforming token-only models by 30% in property prediction. In pocket-based molecular generation, it enhances drug-likeness and synthetic accessibility by approximately 11% and 14%, respectively. Notably, Token-Mol operates 35 times faster than expert diffusion models. In real-world validation, it improves success rates and, when combined with reinforcement learning, further optimizes affinity and drug-likeness, advancing AI-driven drug discovery. In this work the authors present Token-Mol, a token-only 3D drug design model, which deploys the Gaussian cross-entropy (GCE) loss function for regression tasks. It exhibits superior performance in molecular conformation generation, property prediction, and pocket-based generation, thus opening up new avenues for drug design.

Jike Wang, Rui Qin, Mingyang Wang et al. · 30 citations · ⚡1

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.