Skip to content
Preprint

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

Jul 2026 · 0 citations · 26 references
Mathematics Computer Science

TL;DR

A hierarchical Bayesian framework for LLM reliability assessment is extended by relaxing the assumption of independent task outcomes and introducing a Hidden Markov Model to capture sequential dependence in benchmark-constructed interaction sessions, suggesting that ignoring sequential dependence may lead to overconfident reliability estimates.

Abstract

Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, provides a point estimate of performance but does not characterize the uncertainty associated with reliability claims. Currently, statistical inference methods for LLM reliability assessment are emerging. However, a key assumption underlying these models is that test outcomes can be treated as independent repeated trials. This assumption may be inappropriate in sequential settings, where later responses depend on earlier interactions through retained context, error propagation, or an evolving interaction state. We extend a hierarchical Bayesian framework for LLM reliability assessment by relaxing the assumption of independent task outcomes and introducing a Hidden Markov Model to capture sequential dependence in benchmark-constructed interaction sessions. In this formulation, outcomes are generated from a latent interaction state evolving according to a first-order Markov process, capturing changes in interaction context. Through experiments using Anthropic Claude and OpenAI on four datasets, we demonstrate the potential impact of sequential dependence on reliability assessment. The results suggest that ignoring sequential dependence may lead to overconfident reliability estimates.

View source

Similar papers

Preprint Jul 2026

The Computational Basis of Confidence in Large Language Models

Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language models. Existing work has largely evaluated confidence by how well it predicts correctness and whether it is calibrated, leaving open a more fundamental question: what does the confidence signal itself represent? Answer logits may reflect a latent decision variable sufficient to compute normative confidence, or instead a heuristic preference signal that combines the available evidence in a non-Bayesian manner. We address this using statistical decision confidence (SDC), a normative framework from computational neuroscience. Treating the answer-logit difference (LD) as a candidate readout of the latent decision variable, we test the qualitative signatures predicted by SDC. Across three perceptual discrimination tasks and a memory-based decision task, spanning three multimodal non-reasoning models and one reasoning model, LD satisfied these signatures -- including the diagnostic correct/error folded-X pattern -- showing that, in these settings, answer logits behave as monotonic readouts of a latent decision variable rather than heuristic preference scores. In complex visual reasoning, LD continued to predict correctness beyond objective task difficulty, but the full geometric signatures of SDC were absent, illustrating the current boundary of the framework when explicit normative process models are unavailable. These results provide a computational account of confidence in multimodal language models, delineate when answer logits behave as readouts of a latent decision variable, and establish SDC as a unifying framework for studying confidence across biological and artificial intelligence.

D. Kumaran, Viorica Patraucean, M. Ovsjanikov et al. · 0 citations
Open access Jul 2026

Sequential predictive e-diagnostics for hidden Markov models of animal movement

Hidden Markov models are standard for inferring behavioural states from animal movement data, but checking whether a fitted latent-state model predicts held-out movement well remains difficult. We develop sequential predictive e-diagnostics that evaluate a fitted movement HMM as a generator of validation trajectories. Each diagnostic specifies a predictable alternative density, and its ratio to the fitted model’s observable one-step predictive density defines an e-value increment. The denominator is obtained by filtering over latent states, not by conditioning on a decoded path. Under a fixed train/validation protocol, the cumulative product is an e-process, giving anytime-valid thresholds under optional stopping and predictable switching. The construction extends to weighted and state-localized evidence, feature-level circular-linear checks, and blockwise summaries. Controlled simulations show calibration under the fitted-generator null and sensitivity to targeted misspecifications. A leave-one-animal-out elk case study illustrates pooled, individual-specific and state-localized predictive model criticism in a standard movement-HMM workflow.

Aurélien Nicosia · 0 citations
Preprint Jul 2026

BayesAME: Bayesian Active Model Evaluation

Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner to input a coreset size. However, when reliable performance estimation takes priority over efficiency, an evaluation method should also be capable of automatically determining a coreset size that reflects this priority. We introduce BayesAME, a sequential Bayesian framework specifically targeting automatic determination of the coreset size. BayesAME models performance as a random variable by defining a latent ability for each group of items sharing the same historical model performances, with a joint prior distribution encoding the belief that the target model behaves similarly to these historical models. The posterior distribution over these abilities is used to derive performance estimators, quantify performance uncertainty, and select items to add to the coreset via an information-gain criterion. The coreset is iteratively augmented until the performance estimate fluctuation and the performance uncertainty fall below their respective user-defined thresholds. We propose a multi-target extension that captures performance correlations across multiple target models to further reduce the coreset size. Through extensive experiments across diverse benchmarks, we demonstrate that BayesAME consistently outperforms sequential adaptations of existing methods. Crucially, our comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection. Finally, we highlight that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.

Paula Cordero Encinar, taylan. cemgil, Arnaud Doucet et al. · 0 citations
Preprint Jul 2026

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

It is suggested that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates, and this gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.

Patrik Wolf, Thomas Kleine Buening, Andreas Krause et al. · 0 citations
Preprint Jul 2026

Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models

As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable behavioural regularities. Human decision-making combines relatively persistent risk preferences with context-dependent adjustment, yet it remains unclear whether analogous behavioural structure can be observed in LLM-based decision systems. Here we examine this question using a controlled multi-model framework based on no-limit Texas Hold'em, where behaviour is quantified by Participation, measuring voluntary engagement in uncertain opportunities, and Proactiveness, measuring pre-flop risk escalation. Across homogeneous self-play and heterogeneous mixed-model interactions, frontier LLMs exhibit stable, model-specific risk profiles, forming a spectrum from conservative to aggressive decision styles. These profiles remain largely robust under changing opponent composition, while the most conservative and most aggressive models diverge further in mixed settings. Under global risk pressure and personal resource constraint, models adapt in structured but heterogeneous ways, ranging from broad behavioural contraction to selective de-escalation and near-invariant behaviour. These findings suggest that LLMs differ not only in baseline risk disposition, but also in the risk signals they respond to and the flexibility with which they adjust, providing a behavioural basis for auditing risk-sensitive decision-making in interactive settings. Our code is publicly available at: https://github.com/XuankunRong/AgentTexasPoker.

Xuankun Rong, Wenke Huang, Bo Du et al. · 0 citations
Preprint Aug 2026

Integrating Network Psychometrics and LLMs: The Ising-Embeddings-Model applied to Reliability Auditing

Scoring consistency for constructed-response items in large-scale assessments is typically estimated through double-scoring, which uses small samples and assumes independence among responses. We present an integrated framework combining network psychometrics with the Linguistic-Integrated Reliability Audit (LiRA) via a modified Ising model. The model defines a joint distribution over binary correctness labels with pairwise interactions set to the cosine similarity of sentence embeddings and a global bias parameter for item difficulty. LiRA's weighted majority voting over semantic neighborhoods is shown to approximate the conditional logistic distributions of this Ising model. The parametric framework supports benchmark score generation, uncertainty quantification, and missing label imputation; parameters are estimated by maximum pseudo-likelihood. The approach uses the full dataset without requiring extensive double-scoring, accounts for semantic dependencies, and provides diagnostics for rater inconsistencies. The integration of LiRA's scalable methodology with a probabilistic graphical model offers a comprehensive tool for reliability assessment in international assessments such as PIRLS, PISA, and TIMSS.

Matthias von Davier · 0 citations