Skip to content
#small language model Open access

Intent Drift in LLM-Assisted Brain Computer Interface Communication: An In-Silico Benchmark Under Simulated Decoder Corruption

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Language-model post-editing produced fluent semantic substitutions that rose with corruption, confidence did not reliably flag, and no interface policy removed, and this does not demonstrate clinical harm; prospective human-in-the-loop evaluation is needed.

Abstract

Background Large language models are increasingly proposed to post-edit decoded text in communication brain-computer interfaces and augmentative communication. A fluent model can substitute a different intent than attempted (intent drift). Whether meaning survives or confidence flags failure is unmeasured. Methods In-silico benchmark of 20 open-weight models post-editing text (4,252,326 labeled generations) corrupted with an empirical P300 confusion matrix at five levels (0-40% character error rate, CER) across the ALS message-banking vocabulary (AUTH), a message-critical probe set, and matched controls. Outputs were scored faithful, degraded, or drift by an ensemble benchmarked against physicians. A substudy re-ran 562 messages under six interface policies (seven-model panel). Findings Detected drift rose steeply with corruption in all three corpora, from 2.2% to 60.3% at 0-40% target CER in AUTH, a stress-test upper bound, not an expected clinical rate (odds ratio 2.30 per 10-percentage-point rise in target CER). Stated confidence discriminated faithful outputs reasonably well (AUROC 0.83, 0.80-0.85) but was poorly calibrated (expected calibration error 0.32, 0.27-0.37): 28.4% of outputs at confidence 90 or higher were not faithful. Message-critical content carried a small excess after matching, surviving detector removal (rule-free OR 1.10). The ratio of faithful rescues to fluent errors exceeded 1 at low corruption but fell below 1 at 20-30% target CER. No interface policy removed drift: conservative editing and abstention lowered it, alternatives and expansion raised it; the best drifted on 18.0 per 100. A 2,281-item panel (16 of 20 models) gave moderate ensemble-versus-consensus agreement (kappa 0.41); correction lowered pooled drift 31.4% to 28.3%, and a CER-stratified physician-corrected re-analysis confirmed the dose-response at each level. Interpretation Language-model post-editing produced fluent semantic substitutions that rose with corruption, confidence did not reliably flag, and no interface policy removed. This does not demonstrate clinical harm; prospective human-in-the-loop evaluation is needed. Funding: A.G. and E.K. were supported in part by the Clinical and Translational Science Awards (CTSA) grant UL1TR002541 from the National Center for Advancing Translational Sciences, through the Harvard Catalyst | The Harvard Clinical and Translational Science Center Pilot Award Program. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health. Competing interests: The authors declare that they have no competing interests.

Read PDF

Similar papers

Preprint Aug 2026

Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models

Autoregressive large language models (LLMs) routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows. Existing output-stage uncertainty metrics can fail when models are overconfident on false assertions, while multi-sample verification pipelines introduce substantial memory and latency overhead. This work evaluates whether internal hidden-state transition dynamics during generation can signal factual errors without auxiliary decoding calls. We introduce Prediction of Prediction (PoP), a mechanism that captures layer-transition uncertainty by fusing intermediate hidden representations across depth during a single forward pass. Evaluated on the TruthfulQA benchmark using autoregressive transformer backbones, PoP achieves an area under the receiver operating characteristic curve (AUROC) of 75.5% for factual-correctness classification. The mechanism operates within the base forward pass, adding less than 1.2% runtime latency and requiring zero additional generation passes. The numerical results are reported from the author-verified experimental implementation and are bounded by the evaluation scope described below.

Himal Badu · 0 citations
Preprint Aug 2026

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. We present an automatically graded behavioral audit and pilot study of six instruction-tuned VLMs (five general-purpose and one medical specialist) on 4,102 images (4,032 axial/coronal/sagittal MRI slices from 250 subjects plus 70 non-brain/noise controls), with labels derived from public metadata and released expert segmentation masks rather than new human annotation. Across models, answer coverage is near-complete, but verbalized-confidence calibration is poor: ECE ranges from 0.27 to 0.40, mean confidence on incorrect answers ranges from 0.82 to 0.97, and 33-46% of answered items are high-confidence errors. The most accurate model is also the most confident on its errors, while a base/specialist family contrast suggests that medical adaptation improves tumor-presence detection without improving confidence reliability. Open-ended diagnostics further show that hallucination and abstention vary separately from multiple-choice accuracy. These findings argue that medical-image VLM evaluation should report verbalized-confidence reliability, confident error, hallucination, and abstention alongside accuracy.

Amir Sabbaghziarani, Mohammadsajad Abavisani, Sergey M. Plis · 0 citations
Open access Jul 2026

Adaptive Diffusion Vision-Language Models for Reliable Medical Image Understanding

Biomedical vision–language models increasingly support image-grounded clinical dialogue, yet most deployable systems still depend on autoregressive language generation. Such systems tend to truncate answers, react poorly to length instructions, and offer no principled way to signal uncertainty when image evidence is weak. We present MedDiffVL, a biomedical vision-language model that pairs a masked language diffusion backbone with a SigLIP-2 visual encoder and a multimodal alignment pipeline that injects modality and question-type cues. Three inference-time mechanisms target the failure modes of diffusion-based generators in the clinical setting. An adaptive confidence-guided remasking rule uses a time-aware threshold and a short-window stability check to remove repetitive low-quality candidates. A clinically aware length controller selects a target length from question type, modality, and an internal uncertainty estimate. A reliability gate combines visual-evidence and answer-confidence scores to emit, hedge, or escalate a response. On VQA-RAD, SLAKE, and PathVQA, the model reaches 85.42, 92.78, and 94.91% closed-form accuracy and an overall conversation score of 53.42 against a fixed reference. Token repetition falls from 0.18 to 0.06. An ECE falls from 0.137 to 0.034, but this reflects an ECE-surrogate training loss and is not independently validated. These gains are not uniform. The closed-form gains over the prior diffusion model lie within run-to-run variance, and latency stays higher than autoregressive baselines. The main contribution is controllability and reliability-aware decoding, not higher closed-form accuracy. The results indicate that confidence-guided masked diffusion with reliability-aware decoding is a useful direction for controllable and reliability-aware clinical assistants.

Saqib Qamar, Goram Mufarah M. Alshmrani · 0 citations
Preprint Jul 2026

The Capacity of Thought: Benchmarking Llama 3.2 in Semantic fMRI Neural Language Decoding and Improving the Huth Encoding-Model Baseline

Decoding continuous language from fMRI signals remains a core challenge in non-invasive brain-computer interface research. We present two complementary investigations. First, we improve the Huth et al. ridge regression encoding pipeline through expanded voxel selection (10K->15K), substitution of GPT-2 medium for GPT-1 as the beam-search proposal model, and GPU-accelerated bootstrap training, achieving mean METEOR = 0.149 and BLEU-1 = 0.200 across three held-out narratives for subject UTS03 -- an 11% relative METEOR gain over our replication baseline. Second, we introduce fMRIFlamingo, which maps BOLD activity to a frozen Llama-3.2-1B with trainable gated cross-attention layers via a learned brain tokenizer and a Perceiver Resampler. Despite achieving 42.86% Top-1 accuracy on a 1-in-100 ranking task, well above chance, a blind control ablation with zeroed fMRI inputs yields near-identical scores, revealing that apparent decoding success is driven primarily by the frozen language prior rather than by neural input. These results demonstrate that high-capacity language models do not inherently improve fMRI decoding and can actively obscure failures without rigorous blind-control evaluation.

Miško šuvaković, Dom Marhoefer, Glenn Grant-Richards et al. · 0 citations

Related blog posts