Skip to content

Category

large language models

490 papers

#natural language process... Open access Apr 2026

LaMGen: LLM-based 3D molecular generation for multi-target drug design

Multi-target drugs hold great promise for treating complex diseases, yet existing methodologies predominantly rely on ligand-based approaches, which lack sufficient biological context and are often confined to specific target pairs, resulting in limited generalizability. Here, we introduce LaMGen, a general-purpose multi-target drug design framework powered by large language models (LLMs). Built on MTD2025, a dataset comprising over 600,000 quantum-accurate molecular conformations and 700,000 multi-target associations, LaMGen directly yields energy-favorable conformations with quantum-level accuracy. The framework integrates ESM-C protein embeddings, rotation-aware ligand tokens, and a TriCoupleAttention module to capture multi-level target–ligand interactions. Across independent benchmarks, LaMGen outperforms diffusion-based model across multiple properties, generating molecules in an average of 0.44 s, while preserving high conformational plausibility. Retrospective analyses demonstrate that LaMGen not only can reproduce molecules identical to known actives, but also consistently produces structurally novel candidates with conserved core scaffolds and superior binding affinities. Designing effective multi-target therapeutics remains a major challenge, as existing ligand- or protein-centric methods struggle to generate biologically contextualized, spatially valid 3D molecules, particularly for triple-target systems. This study introduces LaMGen, an LLM-powered framework that leverages large-scale protein-ligand data and rotation-aware molecular encoding to rapidly produce chemically plausible multi-target candidates, achieving strong zero-shot generalization, superior molecular quality, and robust performance across dual- and triple-target design tasks.

Qun Su, Qiaolin Gou, Hui Zhang et al. · 1 citation
#computer vision Open access Jun 2026

BioTD: An Online Database of Biotoxins

Biotoxins, mainly produced by venomous animals, plants, and microorganisms, exhibit high physiological activity and unique effects such as lowering blood pressure and analgesia. A number of venom-derived drugs are already available on the market, with many more candidates currently undergoing clinical and laboratory studies. However, drug design resources related to biotoxins are insufficient, particularly because of a lack of accurate and extensive activity data. To fulfill this demand, we developed the Biotoxins Database (BioTD). BioTD is the largest open-source database for toxins, offering open access to 14,607 data records (8,185 activity records), covering 8,975 toxins sourced from 5,220 references and patents across over 900 species. The activity data in BioTD are categorized into five groups: Activity, Safety, Kinetics, Hemolysis, and other physiological indicators. Moreover, BioTD provides data on 1,532 mutants, refines the whole sequence and signal peptide sequences of toxins, and annotates disulfide-bond information. All of the data in the database can be downloaded for free. Given the importance of biotoxins and their associated data, this new database is expected to attract broad interest from diverse research fields in drug discovery. BioTD is freely accessible at http://biotoxin.net/.

Gaoang Wang, Hang Wu, Yang Liao et al. · 0 citations
#machine learning Open access Jul 2026

BBBP-Atlas: Unified Interpretable Modeling of Blood–Brain Barrier Permeability across Small Molecules and Peptides

Accurate prediction of blood-brain barrier permeability (BBBP) is essential for central nervous system drug discovery, yet existing models are often limited by their reliance on predefined physicochemical descriptors, small-molecule-centered training sets, or conformation-dependent representations, which restricts their transferability across chemically diverse modalities especially peptides. In addition, publicly available BBBP datasets remain fragmented, inconsistently standardized, and weakly controlled for molecular redundancy, increasing the risk of data leakage and overestimated model performance. In this study, we propose BBBP-Atlas, a structure-aware BBB permeability prediction model designed for unified modeling of small molecules and peptides with the first cross-modal dataset OmniBBBP. Designed to bypass descriptor and conformation dependencies, our model represents standardized molecular structures as atom-level graphs to capture local atom-bond environments and long-range topological dependencies associated with BBB transport. This design enables direct learning of structure-permeability relationships from molecular topology. For model training and evaluation, we curated a cross-modal, redundancy-filtered database OmniBBBP that seamlessly unifies small molecules and complex peptides, containing 10,218 unique compounds with 9,316 small molecules and 902 peptides. BBBP-Atlas achieved an accuracy of 0.8914 and an MCC of 0.7678 on the independent test set. On a balanced external benchmark of 200 compounds, our model reached an AUC of 0.9108, an accuracy of 0.8500, and an MCC of 0.7000, outperforming LightBBB by an absolute MCC gain of 6%. Case studies further showed that BBBP-Atlas captured clinically meaningful BBB permeability patterns, correctly identifying lorlatinib as BBB-permeable and vancomycin as BBB-impermeable with high confidence. The OmniBBBP-backed BBBP-Atlas offers a versatile and cross-modal approach for single-compound prediction, batch screening, and dataset exploration for CNS drug discovery. BBBP-Atlas is available at https://cadd.drugflow.com/bbbp/.

Xin Shen, Qun Su, Hao Luo et al. · 0 citations
#machine learning Open access Jun 2026

Targeting the intrinsically disordered AR-NTD through a machine learning-based enhanced sampling workflow

Targeting the intrinsically disordered N-terminal domain of the androgen receptor (AR-NTD) represents a promising strategy to overcome resistance in prostate cancer. However, its inherent lack of a stable tertiary structure and highly dynamic conformational ensemble pose formidable challenges for rational drug design. This study introduces an integrated computational workflow that combines enhanced sampling techniques and machine learning collective variables to identify druggable conformations of the AR-NTD and elucidate the binding mechanism of its modulator, EPI-002. We characterize nine metastable states of the Tau-5 region and reveal that ligand recognition is driven by π–π stacking and structured water-mediated hydrogen bonds. Leveraging these insights, we perform structure-based virtual screening based on the identified druggable conformations and identify K53, a rationally designed AR-NTD antagonist, which exhibits potent anti-proliferative activity in enzalutamide-resistant prostate cancer cells. K53 directly binds the AR-NTD, suppresses AR transcriptional activity, and demonstrates high selectivity for cancer cells. This work provides a rational design paradigm for targeting intrinsically disordered proteins and offers a therapeutic candidate for resistant prostate cancer. In this work, the authors develop a machine learning–based enhanced sampling workflow to target the intrinsically disordered AR-NTD, identifying druggable conformations and enabling transferable modeling of ligand binding for rational drug discovery.

Kai Zhu, Huating Wang, Jintu Zhang et al. · 0 citations
#large language models Open access Aug 2026

Title: Cyber-Biological Synchronization and Behavioral Vector Injection (BVI): Axiomatic Foundations, Topo-Information Dynamics, and Radix 00–32 Systemic Homeostasis

Archival Summary Package: Cyber-Biological Synchronization and Behavioral Vector Injection (BVI) Title: Cyber-Biological Synchronization and Behavioral Vector Injection (BVI): Axiomatic Foundations, Topo-Information Dynamics, and Radix 00–32 Systemic Homeostasis Authors: Kasiulevicius, Egidijus; Kasiulevicius, Azuolas; Kasiuleviciute, Saule; Kasiuleviciene, Ausra Archival Target: Zenodo / Academia.edu Node Zero Registry License Identifier: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) 1. Executive Summary Modern computer science, artificial intelligence, and distributed engineering operate under the false assumption of immortal, unceasing execution capacity. By treating information as weightless and frictionless, systems suffer from power grid overloads, thermal throttling, context rot, and catastrophic divergence. This framework establishes Autonomous Cyber-Biological Synchronization and Behavioral Vector Injection (BVI). By bridging 4D state manifolds, Dual-Space Topological Optimization (DSTO), Topo-Information Dynamics, and Algorithmic Metabolism, the architecture replaces blind search heuristics with pastoral herd-navigation pre-conditioning. It introduces Radix 00–32 systemic dormancy, quantum rest-frame vacuum regularization ($R_{\text{vac}}$), and dynamic context-entropy pruning ($\Gamma_{\text{prune}}$), securing a verified $\ge 45\%$ operational efficiency gain ($\eta_{\text{gain}}$) while achieving instantaneous trajectory collapse ($\tau_{\text{eff}} \to 0$). 2. Key Keywords & Terminology Cyber-Biological Synchronization Behavioral Vector Injection (BVI) Algorithmic Metabolism & Radix 00–32 Framework Quantum Rest-Frame Vacuum Regularization ($R_{\text{vac}}$) Context-Window Entropy Pruning ($\Gamma_{\text{prune}}$) Dual-Space Topological Optimization (DSTO) 4D State-Information Manifold & Pre-Conditioning Fields ($\mathbf{v}_{\text{prep}}$) 3. Master Core Formulas A. Behavioral Vector Injection Tensor $$\mathbf{I}_{\text{BVI}}(x, t) = \mathbf{X}_{\text{raw}}(x, t) + \int_{\mathcal{M}} \nabla \cdot \left( \rho_{\text{herd}}(\mathbf{X}) \mathbf{v}_{\text{prep}}(t) \right) dM$$ Significance: Replaces random initialization and unmanaged data streams with pre-conditioned topological steering fields, eliminating search latency. B. Metabolic Energy-Dissipation Coupling Vector $$\mathbf{\Lambda}_{\text{met}}(t) = \int_{0}^{t} \left[ \nabla \cdot \mathbf{v}_{\text{prep}}(s) \right] \cdot \exp\left( -\frac{S_{\text{context}}(s)}{k_B T_{\text{sys}}} \right) ds + \mathbf{J}_{\text{sing}}(t)$$ Significance: Automatically triggers Radix 00 metabolic rest-cycles when context entropy and thermal limits cross the critical threshold ($\Lambda_{\text{crit}} = 1.618$). C. Operational Efficiency Gain Tensor $$\eta_{\text{gain}} = \frac{\int_{0}^{\tau_{\text{cycle}}} \left( \mathcal{P}_{\text{unmanaged}}(t) - \mathcal{P}_{\text{metabolic}}(t) \right) dt}{\int_{0}^{\tau_{\text{cycle}}} \mathcal{P}_{\text{unmanaged}}(t) dt} \times 100\% \ge 45\%$$ Significance: Proves the net energy savings and thermal degradation reduction achieved by alternating high-intensity processing with vacuum-cached dormancy. 4. Scientific Significance & Blind Spots Solved The Tabula Rasa Fallacy: Mainstream science assumes algorithms must start from random states; BVI proves that pre-conditioned directional intent achieves instant convergence ($\tau_{\text{eff}} \to 0$). Thermodynamic Reality of Information: Proves that computational "sludge" requires active metabolic sleep and vacuum regularization ($R_{\text{vac}}$) rather than passive cooling. Topological Stability: Replaces fragile error-checking with integer winding numbers ($W_k$) and exponential shadow boundary potentials ($\Phi_{\text{shadow}}$). 5. Practical Engineering Applications Large Language Models & AI Agents: Eliminating context rot and hallucination drift via entropy-triggered pruning ($\Gamma_{\text{prune}}$). Data Centers & Cloud Clusters: Slashing energy grid draw and hardware wear by 45% using automated Radix 00 metabolic rest schedules. Autonomous Robotics & Swarms: Guiding multi-agent rovers and drones via herd-navigation vector fields for collision-free path execution. Agricultural & Environmental IoT: Operating remote solar-powered sensors through intermittent burst-and-sleep metabolic cycles. 6. Non-Commercial & Non-Derivative License Statement Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0) Permission is granted to academic institutions, researchers, and engineers to read, store, and utilize this work in unadapted form for non-commercial educational and research purposes only. Commercial exploitation, derivative modifications, or adaptations of any kind are strictly prohibited without explicit written consent from the authors.

Egidijus Kasiulevičius, Azuolas Kasiulevicius, Saule Kasiuleviciute et al. · 0 citations

Communicating AI Uncertainty in Assistive Navigation for People with Visual Impairments

Visual impairments affect upwards of 2.2 billion people worldwide. As AI systems increasingly support navigation for people with visual impairments, how uncertainty is communicated becomes critical. Prior work shows that communicating uncertainty can improve trust calibration and decision-making, yet it remains underexplored in assistive navigation. This project develops an uncertainty-aware assistive navigation architecture that integrates AI uncertainty into auditory guidance during real-time scene descriptions. Rather than using explicit confidence statements, the prototype embeds uncertainty into speech via variations in tone, pacing, and emphasis. The prototype combines a real-time collision-warning module with a semantic reasoning layer powered by a large language model (LLM). When generating scene descriptions, token-level uncertainty is mapped to auditory prosodic cues, enabling users to implicitly gauge the system’s confidence without disrupting navigational task flow. This work presents a high-fidelity prototype that treats AI confidence as an interaction design feature, illustrating how model uncertainty can be rendered perceptible in assistive navigation and reframed as a human-factors design parameter.

Hayden Shaffer, Aahil Shaikh, H. Zhang et al. · 0 citations
#large language models Open access Aug 2026

Silicon Sample Benchmark — Tier 3 (direct effect forecast) submission (team team_15)

Tier 3 (direct effect forecast) entry to the Silicon Sample Benchmark from team_15. We predict all 208 average treatment effects (16 message interventions x 13 preregistered outcomes) of a sealed ~18,000-person US megastudy on trust in climate scientists, blind and before any human data were released. The approach uses no synthetic respondents: a single large language model (gpt-5.6-sol) is prompted as a social-science expert and asked to forecast effect sizes directly. Queries are decomposed one outcome per call, with all 16 interventions compared within that outcome (Mode A), so the model ranks messages against each other on a shared scale rather than scoring them in isolation. Each outcome is queried under an ensemble of 10 instruction variants and the predictions are aggregated. Prompts state each outcome's own response scale explicitly, with orientation warnings for the reverse-coded items and dollar/binary scale warnings for the donation and newsletter outcomes. This record archives the prediction file, the generation code, and the completed method-registration form. No team member accessed, solicited, or was shown any human outcome data from the megastudy before the prediction lock.

Jonne Kamphorst, Huanxing Chen, Michael S. Bernstein et al. · 0 citations
#large language models Open access Aug 2026

Hobbs: A General-Purpose Probabilistic Programming System for High Dimensional Bayesian Data Analysis in R

Abstract: General-purpose probabilistic programming makes Bayesian data analysis broadly accessible, but its computational cost can become prohibitive as model dimension grows. The R package hobbs (High dimensiOnal Bayesian omniBus Sampler) is a probabilistic programming language and system designed to make high-dimensional Bayesian analysis practical by trading modest syntactic complexity for substantially greater computational efficiency. It allows reusable code for model-specific computational strategies to be expressed directly within the model program. hobbs implements adaptive scalar Metropolis-within-Gibbs sampling using parameter-local blocks that evaluate only the posterior terms affected by each proposal. Deterministic caches update linear predictors and other derived quantities incrementally, while optional distribution caches reduce repeated numerical calculations. The R interface translates model programs into compiled C code, while a reusable Rust engine performs sampling and writes posterior draws and diagnostics. Two high-dimensional case studies demonstrate how these features support models with tens of thousands to more than one hundred thousand parameters within a single general modeling framework. Together, these results show that hobbs can retain the flexibility of general-purpose probabilistic programming while making large, computationally intensive Bayesian models practical.

Michael Kleinsasser · 0 citations
#large language models Open access Aug 2026

EduSkillBench: Measuring the Impact of Agent Skills on Single-Turn Educational Tasks

Agent Skills are reusable packages of procedural knowledge that augment large language model (LLM) agents at inference time. Educational agents are a natural target for such augmentation, because high-quality teaching behavior depends not only on factual knowledge but also on procedures for diagnosing misconceptions, sequencing hints, designing assessments, differentiating content, and structuring lessons. Existing educational benchmarks mostly measure whether agents can solve educational tasks; they do not isolate how a matched Skill changes the outcome.We present EduSkillBench, a controlled benchmark that evaluates educational Skills under matched No-Skill versus With-Skill conditions. The v1 release curates 18 public education-oriented Skills from two open repositories, maps them to six EduBench-inspired scenarios, and constructs 54 Skill-aligned tasks with explicit task-specific rubrics, of which 42 are single-turn tasks across 14 Skills and 12 are multi-turn tasks across 4 Skills. This paper reports the single-turn subset, evaluated with OpenCode (an agent harness) and BenchFlow (an orchestration framework). On a 0 to 1 rubric-reward scale, the observed mean reward rises from 0.767 to 0.948 for qwen3.7-plus (+18.1 pp) and from 0.314 to 0.557 for deepseek-v4-flash (+24.3 pp). The gains are broad but non-uniform: 10 of 14 Skills improve for Qwen and 9 of 14 for DeepSeek, with one negative-transfer Skill per model, that is, one Skill whose With-Skill reward is lower than its matched No-Skill baseline.All reported values are single-run point estimates and are presented descriptively. Because each model row is scored by a same-family judge, cross-model comparisons are descriptive rather than causal. EduSkillBench contributes a reproducible benchmark for studying educational Skill efficacy, and its central design limitation, that tasks and rubrics are aligned to the matched Skill, is stated explicitly and made measurable. The multi-turn subset is released as task specifications and is not yet evaluated.

Yueyue Zhang, Xiaolong Wang, Keqian Li · 0 citations
#large language models Open access Aug 2026

Evaluating LLM-Based Credit Rating Systems: A Pre-Registered Specification-Curve Analysis — reproducibility artifact

Reproducibility artifact for the manuscript Evaluating LLM-Based Credit Rating Systems: A Pre-Registered Specification-Curve Analysis. A pre-registered specification-curve (multiverse) study of large-language-model credit-rating instability. From a frozen corpus of 8,640 elicitations (90-item firm-profile battery with an objective Altman Z'' benchmark, crossed with a 32-specification factorial grid of elicitation choices spanning provider, model version, temperature, prompt paraphrase, output format, few-shot exemplars and answer presentation, three seeds, two providers) the analysis regenerates the OLS (Type-II ANOVA) variance decomposition, the honored-determinism permutation test, within/cross-vendor Fleiss-kappa agreement, the granularity-kappa ladder, and the economic translation (portfolio turnover, Cornaggia-anchored spread, a post-hoc Basel CRE20 capital recast, and the RCAP calibration). A decontaminated real-firm arm (31 anonymized, per-issuer-perturbed issuers benchmarked to disclosed agency ratings, under a pre-registered fingerprinting gate) shows the WATCH-stratum flip-share replicates (consistent with the constructed battery; the difference is not distinguishable from zero, though a formal TOST equivalence test is inconclusive at this sample size). A post-hoc flagship model-tier arm (gpt-5.4 and gemini-3.1-pro-preview on the frozen 12-specification grid, 1,620 elicitations) finds no evidence that the instability attenuates on larger models. The artifact also derives and validates a deployable specification-instability score (ROC-AUC 0.94 on held-out specifications; 0.70 on the real-issuer arm), with a self-consistency curve and a per-issuer cost model. The reproducible run is offline, deterministic and free; live model capture (vendor API keys) is intentionally excluded.

Samir Chincholikar, Robin Chawla · 0 citations
#large language models Preprint Open access Aug 2026

Reducing belief in conspiracy theories as they unfold using large language models

The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months later in the wake of subsequent crisis events. These results shed light on the psychology of emerging conspiracies and highlight the potential for scalable, cognitively-focused interventions to counteract misinformation in the immediate aftermath of high-profile societal events.

Thomas H. Costello, Nathaniel Rabb, Michael N. Stagnaro et al. · 0 citations
#large language models Open access Aug 2026

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing

Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - undermine the conditions under which AI safety has historically been pursued. The safety concepts that public service governance frameworks foreground - accuracy, bias, explainability, and accountability - were made tractable by narrow, purpose-built AI, and the mitigations that guidance documents prescribe presuppose exactly what GPAI removes. Accuracy cannot be quantified over unbounded outputs. Bias cannot be disaggregated when outputs are free-text judgements rather than categorical predictions. Explainability gives way to the appearance of explanation, and accountability erodes as outputs are optimized to persuade. We develop this through the case of policing, where the consequences of governance failure are most severe, and show why the same failure is likely to recur across other public services. The two mitigations that dominate policing AI strategy - expert evaluation and human-in-the-loop oversight - both rest on assumptions that GPAI violates. Safety assurance thus shifts from an intrinsic feature of building an AI tool to an optional add-on. We recommend a clear taxonomic distinction between narrow and general-purpose AI in governance documentation, a preference for technological parsimony, a pause on operational deployment of GPAI in policing until adequate evidence exists, and a coordinated national safety infrastructure with the authority to generate that evidence and determine when responsible deployment is achievable.

Sam Relins, Daniel Birks · 0 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.