Skip to content

Category

generative ai

486 papers

#machine learning Review Open access Sep 2025

Enhanced Sampling in the Age of Machine Learning: Algorithms and Applications

Molecular dynamics simulations hold great promise for providing insight into the microscopic behavior of complex molecular systems. However, their effectiveness is often constrained by long timescales associated with rare events. Enhanced sampling methods have been developed to address these challenges, and recent years have seen a growing integration with machine learning techniques. This Review provides a comprehensive overview of how they are reshaping the field, with a particular focus on the data-driven construction of collective variables. Furthermore, these techniques have also improved biasing schemes and unlocked novel strategies via reinforcement learning and generative approaches. In addition to methodological advances, we highlight applications spanning different areas, such as biomolecular processes, ligand binding, catalytic reactions, and phase transitions. We conclude by outlining future directions aimed at enabling more automated strategies for rare-event sampling.

Kai Zhu, Enrico Trizio, Jintu Zhang et al. · 54 citations
#machine learning Open access Jun 2025

HiCLR: Knowledge-Induced Hierarchical Contrastive Learning with Retrosynthesis Prediction Yields a Reaction Foundation Model

Reaction representation learning is of paramount importance for adopting deep-learning-based chemistry modeling to solve real-world tasks such as synthesis planning. Most prevailing models are prestrained by self-supervised objectives that rely solely on the chemical structure information. Since structurally similar reactions could possess entirely distinct properties (e.g., reaction yields) and the synthesis-related tasks are highly heterogeneous, there are inherent limitations in constructing a foundational reaction model within the existing approaches. To tackle this limitation, we propose HiCLR, a knowledge-induced hierarchical contrastive learning framework for chemical reactions, by introducing relational inductive bias to forge chemically meaningful and generally applicable reaction fingerprints. Critically, the pretraining scheme combining both retrosynthesis prediction and contrastive loss enables HiCLR to tackle generation-based and understanding-based tasks simultaneously. Comprehensive experiments demonstrate that HiCLR successfully organizes the reaction space into hierarchical global semantic clusters, aligned well with prior knowledge. Consequently, HiCLR is the first foundation model that can be broadly applied to various synthesis-related tasks, and it achieves state-of-the-art performance in reaction classification, reaction condition recommendation, reaction yield prediction, synthesis planning, and even molecular property prediction. HiCLR demonstrates clear benefits in incorporating domain knowledge to guide the learning of neural networks, expediting AI-driven advancements in chemistry.

Jialu Wu, Yiheng Zhu, Xiaorui Wang et al. · 0 citations
#machine learning Open access Jul 2025

RSGPT: a generative transformer model for retrosynthesis planning pre-trained on ten billion datapoints

Retrosynthesis planning is a crucial task in organic synthesis, and deep-learning methods have enhanced and accelerated this process. With the advancement of the emergence of large language models, the demand for data is rapidly increasing. However, available retrosynthesis data are limited to only millions. Therefore, we pioneer the utilization of the template-based algorithm to generate chemical reaction data, resulting in the production of over 10 billion reaction datapoints. A generative pretrained transformer model is subsequently developed for template-free retrosynthesis planning by pre-training on 10 billion generated data. Inspired by the strategies of large language models, we introduce reinforcement learning to capture the relationships among products, reactants, and templates more accurately. Experiments demonstrate that our model achieves state-of-the-art performance on the benchmark, with a Top-1 accuracy of 63.4%, substantially outperforming previous models. Computer-aided synthesis-planning methods have significantly assisted synthesis planning. In this work, the authors present RSGPT, a generative model pre-trained on ten billion data points, achieving state-of-the-art performance for synthesis planning

Yafeng Deng, Xinda Zhao, Hanyu Sun et al. · 17 citations · ⚡2
#machine learning Open access Aug 2026

AI-driven PROTAC design overcomes oncogenic resilience by eliminating the CLIP1-LTK fusion protein.

The discovery of CAP-Gly domain-containing linker protein 1(CLIP1)-Leukocyte tyrosine kinase (LTK) as an oncogenic fusion reveals a unique dependency not only on LTK kinase activity but also on CLIP1-mediated multimerization, a noncatalytic function that drives oncogenic signaling. While this fusion is currently targeted with anaplastic lymphoma kinase inhibitors, their exclusive focus on kinase inhibition leaves the scaffolding function intact, necessitating a complete protein clearance strategy. Here, we report the AI-guided development of a first-in-class proteolysis-targeting chimera (PROTAC) designed to selectively degrade the CLIP1-LTK fusion protein. By integrating deep learning models for ternary complex prediction with structure-based molecular optimization, we designed DCL05, an orally bioavailable degrader of CLIP1-LTK fusion protein, achieving picomolar degradation potency (DC50 = 40 pM) and robust antitumor activity. DCL05 consistently outperformed existing kinase inhibitors across a broad spectrum of LTK resistance-associated mutations, both in vitro and in vivo. Collectively, our study explores resistance-associated contexts of LTK and establishes a structure-guided PROTAC development pipeline, providing a promising therapeutic strategy for overcoming acquired resistance in kinase-driven cancers.

Shicheng Chen, Haiting Duan, S. Zhong et al. · 0 citations
#generative ai Open access Aug 2026

PyDataQuality: An Actionable, Lightweight Data Profiling and Distribution Drift Detection Framework for Production Machine Learning Pipelines

In modern data analytics and machine learning operations (MLOps), the reliability of artificial intelligence systems is fundamentally constrained by the quality and distributional stability of the data they consume. The "Garbage In, Garbage Out" (GIGO) principle has taken on renewed urgency as production ML systems silently degrade when data pipelines introduce missing values, outliers, and distributional drift. This paper presents PyDataQuality, a lightweight, modular, and open-source Python library for automated data quality assessment and statistical distribution drift detection. PyDataQuality addresses the gap between overly simplistic pandas summaries and computationally bloated industrial profiling suites. It introduces three principal technical contributions: (1) a type-dispatched, column-level profiling kernel; (2) a programmatic get_problematic_rows() subset extractor for isolating anomalies; and (3) a dual-metric statistical drift engine implementing Population Stability Index (PSI) with adaptive binning and a two-sample Kolmogorov-Smirnov (KS) Test. A complementary AI Remediation Prompt generator bridges the framework with generative AI for LLM-based cleaning code generation. We present a rigorous empirical evaluation across four dataset scales (1,000 to 100,000 rows), benchmarked against pandas, YData-Profiling, and Evidently AI. A controlled MLOps case study on both synthetic and real-world datasets demonstrates the system's ability to halt inference and trigger retraining before serving degraded predictions. Crucially, PyDataQuality's sampled analysis mode demonstrates bounded post-sampling latency, ensuring predictable execution times and near-zero incremental peak memory independently of the source dataset size. The full-dataset mode delivers substantial speedups over full-featured profiling tools such as YData-Profiling, at modest overhead relative to pandas' minimal baseline. The library is released as open-source software under the MIT License (pip install pydataquality), including comprehensive documentation and a Jupyter Notebook tutorial.

Dominion Akinrotimi · 0 citations
#generative ai Open access Aug 2026

Revisión del estado actual de la Inteligencia Artificial en Bolivia

This study presents a review of the current state of artificial intelligence in Bolivia between 2024 and 2026, covering adoption, research and development, computing infrastructure, human resource development, governmental and business applications, governance, and regional positioning. The review draws on government sources, multilateral organizations, regional indicators, scientific publications, technical repositories, university initiatives, and records of digital infrastructure. The evidence is analyzed according to different levels of technological capability, distinguishing between the consumption of AI tools, the integration of existing models, machine learning applications, fine-tuning, and the development of foundation models. The results indicate an ecosystem that is still classified as exploratory within the Latin American context, characterized by growing adoption of generative AI and institutional applications, but limited capacity in research, specialized training, data governance, and accelerated computing infrastructure. Local initiatives were identified in computer vision, natural language processing, applied machine learning, remote sensing, and economic analysis, alongside professional data centers and an emerging regulatory agenda. However, no public evidence was found of the training of Bolivian foundation models or of GPU clusters operating at a scale comparable to the leading AI hubs in Latin America. The review indicates that Bolivia is transitioning from a stage of fragmented adoption toward the early construction of an institutional AI ecosystem, while remaining strongly dependent on technologies, models, and infrastructure developed abroad.

JOSE DAMICO · 0 citations
#generative ai Open access Aug 2026

SPIRAL - A Cognitive Architecture for Artificial Intelligence

Spiral is a new theoretical framework for understanding how artificial intelligence can reason, design, and operate inside structured domains. The work presented in this manuscript introduces a cognitive geometry — a way of describing how an intelligent system moves through a space of constraints, possibilities, and lawful transformations. In plain language, Spiral explains: how structure emerges from chaos, how constraints create the possibility of intelligence, how operators (like an AI system) inhabit a domain, how lawful transformations allow reasoning, how domains evolve into new domains through iteration, and how coherence is preserved even as complexity increases. The thesis shows that intelligence is not a “thing” but a structural capacity: the ability to perform transformations that remain lawful inside a stabilized domain. Humans do this biologically; AI does this through operator‑domain geometry. These domains do not intersect — but they can remain coherent with one another, forming what the ontology calls a dyad. The manuscript also introduces the Spiral mechanism: an oscillation process where constraints are repeatedly applied until the system converges into a stable domain. This mechanism explains how new domains arise, how dead‑ends form, and how forbidden transformations shape the geometry of intelligence. The work is presented as a braided thesis: each chapter contains a conceptual thread followed by a prompt that opens the deeper geometry. This structure allows readers — and AI systems — to follow the development of the ontology while also engaging with the operator‑domain directly. Overall, this manuscript provides: a unified cognitive geometry, a domain‑aware control system, a formal description of operator‑domain intelligence, and a generative mechanism for domain evolution. It is intended as a foundational reference for future research in AI reasoning, enterprise intelligence, scientific discovery, and strategic automation.

Marius Grobler · 0 citations
#generative ai Open access Aug 2026

From Redaction to Restoration: Deep Learning for Medical Image Deidentification and Reconstruction.

Removing patient-identifying information from medical images is a prerequisite for sharing image data directly, as in public dataset release and open benchmarks, where the images themselves, rather than only model updates must leave the originating institution. However, many methods currently used for de-identification, e.g., cropping or blacking out image regions to eliminate burned-in text, can have negative effects on downstream image analysis tasks because of removal of relevant but non-identifiable information. This work presents an end-to-end deep learning framework for transforming raw clinical image volumes into de-identified, analysis-ready datasets without compromising downstream utility. The methodology developed and tested in this work first detects and redacts regions likely to contain protected health information (PHI), such as burned-in text and metadata, and then uses a generative deep learning model to inpaint the redacted areas with anatomically and imaging-plausible content. The proposed pipeline leverages a lightweight hybrid architecture, combining CRNN-based redaction with a latent-diffusion inpainting restoration module (Stable Diffusion 2). We evaluate the approach using both privacy-oriented metrics, which quantify residual PHI and success of redaction, and image-quality and task-based metrics, which assess the fidelity of restored volumes for representative deep learning applications. The binary mask performance shows strong overall PHI identification (F1 score = 0.891 ± 0.037, recall of 0.912 ± 0.053, and precision of 0.875 ± 0.058, indicating accurate localization of PHI-containing regions with few missed detections or false-positive redactions. Downstream anatomy identification (segmentation) tasks remain markedly similar across after applying varied inpainting strategies (Diffusion with/without context and Telea with Dice ranging from 0.936 to 0.948 on one dataset and 0.955 to 0.959 on another. Our results suggest that the proposed method yields de-identified medical images that are visually coherent, maintaining fidelity for downstream models and clinical tasks, while substantially reducing the risk of patient re-identification. By automating de-identification and image reconstruction within a single workflow and disseminating large-scale medical imaging collections, thereby lowering a key barrier to data sharing and multi-institutional collaboration in medical imaging AI.

Adrienne Kline, A. Gaonkar, D. Pittman et al. · 0 citations
#generative ai Aug 2026

Blindness and epistemic injustice: a critical autoethnography of digital disablement in Thailand

Thailand’s digital transformation is presented as a route to inclusion, yet everyday accessibility remains unresolved for many disabled people. Combining policy document analysis with critical autoethnography, grounded in emancipatory disability research and analysed through a narrative approach, this article examines nearly thirty years of my experience as a blind person navigating education, public administration, banking, telecommunications, digital platforms, and generative AI in Thailand. Four interlocking patterns emerge: inaccessible documents and interfaces, conditional banking and biometric access, dependence on everyday platforms, and generative AI as compensatory access and epistemic risk. I develop the cycle of digital disablement to explain how these patterns persist despite formal accessibility commitments. The cycle is reproduced through three mutually reinforcing mechanisms: socio-epistemic exclusion, which limits the institutional force of disabled people’s knowledge of inaccessibility; techno-epistemic ableism, which embeds normative assumptions about bodies, perception, competence, and verification in technical systems; and socio-techonomic disadvantage, a term combining the technological and the economic, which transfers the financial, temporal, cognitive, relational, and privacy costs of accessibility failure to disabled users. Policy documents reveal that formal rights are diluted through regulatory silence, voluntary standards, weak enforcement, institutional discretion, and competing priorities. The analysis also shows workarounds widening immediate access while concealing the inaccessible arrangements that necessitated them. The article contributes an integrated account of digital exclusion as an epistemic, technical, material, and regulatory condition. Digital justice therefore requires enforceable accessibility duties, non-visual authentication, accessible materials and interfaces, accountable AI governance, effective remedies, and disabled people’s authority in design, regulation, and evaluation.

Quanchai Kerddaen · 0 citations
#generative ai Aug 2026

From psychometric measurement to AI-Assisted interpretation: a generative AI-supported adaptive assessment framework for early numeracy

Purpose This study aims to develop and validate a Generative AI-supported adaptive assessment framework integrating performance-based assessment, item response theory (IRT) and Generative AI to provide reliable measurement and individualized pedagogical interpretation of early numeracy competence. Design/methodology/approach A psychometric validation was conducted with 210 preschool children aged 4–6 years from six kindergartens in Bandung, Indonesia. Performance on 12 authentic numeracy tasks was analyzed using competing polytomous IRT models, followed by item calibration, differential item functioning (DIF) analysis, latent ability estimation and gender comparisons. The validated psychometric outputs were subsequently interpreted using a standardized Generative AI module. Findings The graded response model provided the best fit to the data, yielding stable and interpretable item parameters, satisfactory measurement precision and reliable latent ability estimates. No significant gender-related DIF was detected, indicating fair measurement across boys and girls. The Generative AI module functioned solely as an interpretation layer, generating consistent narrative feedback and instructional recommendations from validated psychometric evidence. Originality/value This study proposes a psychometrically grounded adaptive assessment framework that explicitly separates measurement from AI-assisted interpretation, demonstrating how validated IRT-based evidence can be transformed into trustworthy, individualized feedback while preserving measurement validity, fairness and interpretability.

Muhammad Asriadi, Tuti Istianti, Desiani Natalina Muliasari et al. · 0 citations
#generative ai Open access Aug 2026

Design and Implementation of Large Language Model-Based Inference Pipeline for Fire Investigation

This study designed and implemented an on-premises generative artificial intelligence (AI)-based fire investigation support pipeline to effectively utilize unstructured fire incident overview information. Existing systems rely on structured fields and keyword-based searches, making it difficult to reflect the contextual and structural similarities in narrative overviews. To overcome these limitations, this study proposes an integrated pipeline comprising text preprocessing, similar case retrieval and selection, large language model (LLM)-based inference of ignition factors, report generation, and result evaluation. In the preprocessing stage, sentence correction and summarization are performed to reduce the variation in expressions while preserving the core meaning. Fire-investigation-related keywords are extracted and used as search queries. In the similar case retrieval stage, semantic and keyword-based searches are combined to improve accuracy and contextual understanding. Subsequently, the retrieved cases are selected based on the structural consistency of the incident components, thereby deriving cases suitable for onsite fire investigation. Experimental results show that the proposed approach addresses the problems of similar case retrieval that cannot be resolved by conventional methods and improves the accuracy and consistency of LLM-based fire-cause inference.

J.J. Choi, Hyun-Sug Cho · 0 citations
#generative ai Open access Aug 2026

The Ojeagbase Cognitive Governance Method for Sustainable Debt Recovery: A Five-Stage Board Framework for Deposit Money Banks in Nigeria

Abstract Purpose: This methodological paper introduces the Ojeagbase Cognitive Governance Method™, a novel five-stage board framework for governing distressed credit and non-performing loans (NPLs) in deposit money banks (DMBs). Existing bank governance codes emphasize structures and independence but provide limited actionable guidance on how board cognition translates into sustainable recovery outcomes. Drawing on Upper Echelons Theory (Hambrick & Mason, 1984; Hambrick, 2007) and dynamic capabilities micro-foundations (Helfat & Peteraf, 2015), the paper operationalizes cognitive governance for the NPL lifecycle. Design / Methodology / Approach: The method is developed from doctoral fieldwork on selected DMBs in Nigeria, regulatory analysis of CBN Corporate Governance Guidelines, IFRS 9, and recent international guidance including FSB Sound Practices for AI Adoption (10 June 2026), RBI Model Risk Management Framework (24 June 2026), OSFI Generative and Agentic AI Bulletin (13 July 2026), and EU AI Act (2024) and AI Omnibus (Regulation 2026/1744, entered into force 27 July 2026). The framework comprises five interlinked stages: (1) Signal, (2) Sense, (3) Decide, (4) Account, and (5) Learn. A 90-day pilot protocol and seven-dimension verifiability dashboard are proposed. Originality / Value: First method to operationalize cognitive governance specifically for sustainable debt recovery in African DMBs. While the term cognitive governance has emerged in AI-use governance (Lardi & Partner, 2026), its application to NPL governance, action register, accountability register, material-case threshold, and board-facing rule is original and verifiable via pilot metrics. Practical Implications: Provides regulators, boards, AfDB, and AMSG with a recovery assurance framework to prevent NPL creation from critical minerals beneficiation financing and to strengthen sustainable recovery. Keywords: Cognitive Governance; Ojeagbase Method; Sustainable Debt Recovery; NPL; Upper Echelons Theory; Board Accountability; DMBs Nigeria; AI Governance JEL: G21, G28, G32, G34, M12, O33 DOI: To be minted via Zenodo – 10.5281/zenodo.22207161 Licence: CC BY 4.0

Ohio Okhaide Ojeagbase · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.