Multi-target drugs hold great promise for treating complex diseases, yet existing methodologies predominantly rely on ligand-based approaches, which lack sufficient biological context and are often confined to specific target pairs, resulting in limited generalizability. Here, we introduce LaMGen, a general-purpose multi-target drug design framework powered by large language models (LLMs). Built on MTD2025, a dataset comprising over 600,000 quantum-accurate molecular conformations and 700,000 multi-target associations, LaMGen directly yields energy-favorable conformations with quantum-level accuracy. The framework integrates ESM-C protein embeddings, rotation-aware ligand tokens, and a TriCoupleAttention module to capture multi-level target–ligand interactions. Across independent benchmarks, LaMGen outperforms diffusion-based model across multiple properties, generating molecules in an average of 0.44 s, while preserving high conformational plausibility. Retrospective analyses demonstrate that LaMGen not only can reproduce molecules identical to known actives, but also consistently produces structurally novel candidates with conserved core scaffolds and superior binding affinities. Designing effective multi-target therapeutics remains a major challenge, as existing ligand- or protein-centric methods struggle to generate biologically contextualized, spatially valid 3D molecules, particularly for triple-target systems. This study introduces LaMGen, an LLM-powered framework that leverages large-scale protein-ligand data and rotation-aware molecular encoding to rapidly produce chemically plausible multi-target candidates, achieving strong zero-shot generalization, superior molecular quality, and robust performance across dual- and triple-target design tasks.
Qun Su, Qiaolin Gou, Hui Zhang et al.· Nature Communications· 1 citation
Biotoxins, mainly produced by venomous animals, plants, and microorganisms, exhibit high physiological activity and unique effects such as lowering blood pressure and analgesia. A number of venom-derived drugs are already available on the market, with many more candidates currently undergoing clinical and laboratory studies. However, drug design resources related to biotoxins are insufficient, particularly because of a lack of accurate and extensive activity data. To fulfill this demand, we developed the Biotoxins Database (BioTD). BioTD is the largest open-source database for toxins, offering open access to 14,607 data records (8,185 activity records), covering 8,975 toxins sourced from 5,220 references and patents across over 900 species. The activity data in BioTD are categorized into five groups: Activity, Safety, Kinetics, Hemolysis, and other physiological indicators. Moreover, BioTD provides data on 1,532 mutants, refines the whole sequence and signal peptide sequences of toxins, and annotates disulfide-bond information. All of the data in the database can be downloaded for free. Given the importance of biotoxins and their associated data, this new database is expected to attract broad interest from diverse research fields in drug discovery. BioTD is freely accessible at http://biotoxin.net/.
Gaoang Wang, Hang Wu, Yang Liao et al.· Journal of Chemical Informat...· 0 citations
Accurate prediction of blood-brain barrier permeability (BBBP) is essential for central nervous system drug discovery, yet existing models are often limited by their reliance on predefined physicochemical descriptors, small-molecule-centered training sets, or conformation-dependent representations, which restricts their transferability across chemically diverse modalities especially peptides. In addition, publicly available BBBP datasets remain fragmented, inconsistently standardized, and weakly controlled for molecular redundancy, increasing the risk of data leakage and overestimated model performance. In this study, we propose BBBP-Atlas, a structure-aware BBB permeability prediction model designed for unified modeling of small molecules and peptides with the first cross-modal dataset OmniBBBP. Designed to bypass descriptor and conformation dependencies, our model represents standardized molecular structures as atom-level graphs to capture local atom-bond environments and long-range topological dependencies associated with BBB transport. This design enables direct learning of structure-permeability relationships from molecular topology. For model training and evaluation, we curated a cross-modal, redundancy-filtered database OmniBBBP that seamlessly unifies small molecules and complex peptides, containing 10,218 unique compounds with 9,316 small molecules and 902 peptides. BBBP-Atlas achieved an accuracy of 0.8914 and an MCC of 0.7678 on the independent test set. On a balanced external benchmark of 200 compounds, our model reached an AUC of 0.9108, an accuracy of 0.8500, and an MCC of 0.7000, outperforming LightBBB by an absolute MCC gain of 6%. Case studies further showed that BBBP-Atlas captured clinically meaningful BBB permeability patterns, correctly identifying lorlatinib as BBB-permeable and vancomycin as BBB-impermeable with high confidence. The OmniBBBP-backed BBBP-Atlas offers a versatile and cross-modal approach for single-compound prediction, batch screening, and dataset exploration for CNS drug discovery. BBBP-Atlas is available at https://cadd.drugflow.com/bbbp/.
Xin Shen, Qun Su, Hao Luo et al.· bioRxiv· 0 citations
Targeting the intrinsically disordered N-terminal domain of the androgen receptor (AR-NTD) represents a promising strategy to overcome resistance in prostate cancer. However, its inherent lack of a stable tertiary structure and highly dynamic conformational ensemble pose formidable challenges for rational drug design. This study introduces an integrated computational workflow that combines enhanced sampling techniques and machine learning collective variables to identify druggable conformations of the AR-NTD and elucidate the binding mechanism of its modulator, EPI-002. We characterize nine metastable states of the Tau-5 region and reveal that ligand recognition is driven by π–π stacking and structured water-mediated hydrogen bonds. Leveraging these insights, we perform structure-based virtual screening based on the identified druggable conformations and identify K53, a rationally designed AR-NTD antagonist, which exhibits potent anti-proliferative activity in enzalutamide-resistant prostate cancer cells. K53 directly binds the AR-NTD, suppresses AR transcriptional activity, and demonstrates high selectivity for cancer cells. This work provides a rational design paradigm for targeting intrinsically disordered proteins and offers a therapeutic candidate for resistant prostate cancer. In this work, the authors develop a machine learning–based enhanced sampling workflow to target the intrinsically disordered AR-NTD, identifying druggable conformations and enabling transferable modeling of ligand binding for rational drug discovery.
Kai Zhu, Huating Wang, Jintu Zhang et al.· Nature Communications· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Visual impairments affect upwards of 2.2 billion people worldwide. As AI systems increasingly support navigation for people with visual impairments, how uncertainty is communicated becomes critical. Prior work shows that communicating uncertainty can improve trust calibration and decision-making, yet it remains underexplored in assistive navigation. This project develops an uncertainty-aware assistive navigation architecture that integrates AI uncertainty into auditory guidance during real-time scene descriptions. Rather than using explicit confidence statements, the prototype embeds uncertainty into speech via variations in tone, pacing, and emphasis. The prototype combines a real-time collision-warning module with a semantic reasoning layer powered by a large language model (LLM). When generating scene descriptions, token-level uncertainty is mapped to auditory prosodic cues, enabling users to implicitly gauge the system’s confidence without disrupting navigational task flow. This work presents a high-fidelity prototype that treats AI confidence as an interaction design feature, illustrating how model uncertainty can be rendered perceptible in assistive navigation and reframed as a human-factors design parameter.
Hayden Shaffer, Aahil Shaikh, H. Zhang et al.· Proceedings of the Human Fac...· 0 citations
Tier 3 (direct effect forecast) entry to the Silicon Sample Benchmark from team_15. We predict all 208 average treatment effects (16 message interventions x 13 preregistered outcomes) of a sealed ~18,000-person US megastudy on trust in climate scientists, blind and before any human data were released. The approach uses no synthetic respondents: a single large language model (gpt-5.6-sol) is prompted as a social-science expert and asked to forecast effect sizes directly. Queries are decomposed one outcome per call, with all 16 interventions compared within that outcome (Mode A), so the model ranks messages against each other on a shared scale rather than scoring them in isolation. Each outcome is queried under an ensemble of 10 instruction variants and the predictions are aggregated. Prompts state each outcome's own response scale explicitly, with orientation warnings for the reverse-coded items and dollar/binary scale warnings for the donation and newsletter outcomes. This record archives the prediction file, the generation code, and the completed method-registration form. No team member accessed, solicited, or was shown any human outcome data from the megastudy before the prediction lock.
Jonne Kamphorst, Huanxing Chen, Michael S. Bernstein et al.· Zenodo (CERN European Organi...· 0 citations
Abstract: General-purpose probabilistic programming makes Bayesian data analysis broadly accessible, but its computational cost can become prohibitive as model dimension grows. The R package hobbs (High dimensiOnal Bayesian omniBus Sampler) is a probabilistic programming language and system designed to make high-dimensional Bayesian analysis practical by trading modest syntactic complexity for substantially greater computational efficiency. It allows reusable code for model-specific computational strategies to be expressed directly within the model program. hobbs implements adaptive scalar Metropolis-within-Gibbs sampling using parameter-local blocks that evaluate only the posterior terms affected by each proposal. Deterministic caches update linear predictors and other derived quantities incrementally, while optional distribution caches reduce repeated numerical calculations. The R interface translates model programs into compiled C code, while a reusable Rust engine performs sampling and writes posterior draws and diagnostics. Two high-dimensional case studies demonstrate how these features support models with tens of thousands to more than one hundred thousand parameters within a single general modeling framework. Together, these results show that hobbs can retain the flexibility of general-purpose probabilistic programming while making large, computationally intensive Bayesian models practical.
Michael Kleinsasser· Zenodo (CERN European Organi...· 0 citations
Agent Skills are reusable packages of procedural knowledge that augment large language model (LLM) agents at inference time. Educational agents are a natural target for such augmentation, because high-quality teaching behavior depends not only on factual knowledge but also on procedures for diagnosing misconceptions, sequencing hints, designing assessments, differentiating content, and structuring lessons. Existing educational benchmarks mostly measure whether agents can solve educational tasks; they do not isolate how a matched Skill changes the outcome.We present EduSkillBench, a controlled benchmark that evaluates educational Skills under matched No-Skill versus With-Skill conditions. The v1 release curates 18 public education-oriented Skills from two open repositories, maps them to six EduBench-inspired scenarios, and constructs 54 Skill-aligned tasks with explicit task-specific rubrics, of which 42 are single-turn tasks across 14 Skills and 12 are multi-turn tasks across 4 Skills. This paper reports the single-turn subset, evaluated with OpenCode (an agent harness) and BenchFlow (an orchestration framework). On a 0 to 1 rubric-reward scale, the observed mean reward rises from 0.767 to 0.948 for qwen3.7-plus (+18.1 pp) and from 0.314 to 0.557 for deepseek-v4-flash (+24.3 pp). The gains are broad but non-uniform: 10 of 14 Skills improve for Qwen and 9 of 14 for DeepSeek, with one negative-transfer Skill per model, that is, one Skill whose With-Skill reward is lower than its matched No-Skill baseline.All reported values are single-run point estimates and are presented descriptively. Because each model row is scored by a same-family judge, cross-model comparisons are descriptive rather than causal. EduSkillBench contributes a reproducible benchmark for studying educational Skill efficacy, and its central design limitation, that tasks and rubrics are aligned to the matched Skill, is stated explicitly and made measurable. The multi-turn subset is released as task specifications and is not yet evaluated.
Reproducibility artifact for the manuscript Evaluating LLM-Based Credit Rating Systems: A Pre-Registered Specification-Curve Analysis. A pre-registered specification-curve (multiverse) study of large-language-model credit-rating instability. From a frozen corpus of 8,640 elicitations (90-item firm-profile battery with an objective Altman Z'' benchmark, crossed with a 32-specification factorial grid of elicitation choices spanning provider, model version, temperature, prompt paraphrase, output format, few-shot exemplars and answer presentation, three seeds, two providers) the analysis regenerates the OLS (Type-II ANOVA) variance decomposition, the honored-determinism permutation test, within/cross-vendor Fleiss-kappa agreement, the granularity-kappa ladder, and the economic translation (portfolio turnover, Cornaggia-anchored spread, a post-hoc Basel CRE20 capital recast, and the RCAP calibration). A decontaminated real-firm arm (31 anonymized, per-issuer-perturbed issuers benchmarked to disclosed agency ratings, under a pre-registered fingerprinting gate) shows the WATCH-stratum flip-share replicates (consistent with the constructed battery; the difference is not distinguishable from zero, though a formal TOST equivalence test is inconclusive at this sample size). A post-hoc flagship model-tier arm (gpt-5.4 and gemini-3.1-pro-preview on the frozen 12-specification grid, 1,620 elicitations) finds no evidence that the instability attenuates on larger models. The artifact also derives and validates a deployable specification-instability score (ROC-AUC 0.94 on held-out specifications; 0.70 on the real-issuer arm), with a self-consistency curve and a per-issuer cost model. The reproducible run is offline, deterministic and free; live model capture (vendor API keys) is intentionally excluded.
Samir Chincholikar, Robin Chawla· Open MIND· 0 citations
The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months later in the wake of subsequent crisis events. These results shed light on the psychology of emerging conspiracies and highlight the potential for scalable, cognitively-focused interventions to counteract misinformation in the immediate aftermath of high-profile societal events.
Thomas H. Costello, Nathaniel Rabb, Michael N. Stagnaro et al.· 0 citations
Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. That pressure has intensified with general-purpose AI (GPAI): AI built on large language models that can be directed by prompt alone to perform an effectively unbounded range of tasks. We argue that the properties that make these models attractive - their generality, accessibility, and low deployment cost - undermine the conditions under which AI safety has historically been pursued. The safety concepts that public service governance frameworks foreground - accuracy, bias, explainability, and accountability - were made tractable by narrow, purpose-built AI, and the mitigations that guidance documents prescribe presuppose exactly what GPAI removes. Accuracy cannot be quantified over unbounded outputs. Bias cannot be disaggregated when outputs are free-text judgements rather than categorical predictions. Explainability gives way to the appearance of explanation, and accountability erodes as outputs are optimized to persuade. We develop this through the case of policing, where the consequences of governance failure are most severe, and show why the same failure is likely to recur across other public services. The two mitigations that dominate policing AI strategy - expert evaluation and human-in-the-loop oversight - both rest on assumptions that GPAI violates. Safety assurance thus shifts from an intrinsic feature of building an AI tool to an optional add-on. We recommend a clear taxonomic distinction between narrow and general-purpose AI in governance documentation, a preference for technological parsimony, a pause on operational deployment of GPAI in policing until adequate evidence exists, and a coordinated national safety infrastructure with the authority to generate that evidence and determine when responsible deployment is achievable.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.