While large language models (LLMs) are having a transformative impact on human society, evaluating them remains challenging. Standard benchmarks usually rely on single point estimates that obscure response stochasticity and variability in question difficulty. Here, we introduce a hierarchical Bayesian Beta-Binomial framework for uncertainty-aware evaluation of LLMs on multiple-choice datasets. Our approach moves beyond single accuracy metrics by modeling correct responses binomially and decomposing performance variation into intra-question stochasticity (response variability for a given question) and inter-question heterogeneity (variation in difficulty across questions) using separate priors. The framework yields a probabilistic assessment, providing full posterior distributions and credible intervals for mean accuracy, inter-question heterogeneity, and mean intra-question response variability, enabling rigorous uncertainty quantification. We demonstrate its utility by evaluating multiple LLMs across diverse benchmarks, including under semantic perturbations like question rephrasing. This analysis reveals nuanced model robustness insights and uncovers distinct behaviors across model classes (e.g., reasoning vs. non-reasoning) often missed by traditional descriptive statistics. This methodology offers a statistically grounded and powerful Bayesian lens for analyzing LLM performance, providing deeper insights into accuracy, consistency, and heterogeneity essential for reliable model benchmarking.
Giannis Manousaridis, John G. Samuelsson, B. Emir et al.· Frontiers in Applied Mathema...· 0 citations
Productivity in pharmaceutical R&D continues to fall despite deeper biological insight and steady gains in clinical development operations - a phenomenon termed Eroom's Law. Agentic AI workflows powered by reasoning-trained large language models (LLMs), increasingly described as large reasoning models (LRMs), could potentially dent this trend. Unlike earlier task-specific models, these systems couple multi-step reasoning with the ability to plan, invoke external tools and retrieve authoritative information, enabling them to decompose and execute complex scientific and operational tasks. This review discusses agentic AI applications across the drug development continuum, from target discovery to post-market surveillance, and highlights three near-term use cases: algorithmic drug repurposing, informed consent support and automated drafting of regulatory documents. For each, we outline plausible architectures, the current level of supporting evidence and the principal failure modes that constrain deployment. Realizing these gains, however, requires prospective validation, rigorous human oversight and governance frameworks that align algorithmic outputs with clinical, ethical, legal and regulatory standards. When implemented responsibly, agentic AI could transform human-AI collaboration in biopharma, improving R&D efficiency and accelerating delivery of safer, more-effective therapies.
Sheraz Khan, John G. Samuelsson, X. Cai et al.· Drug Discovery Today· 0 citations