Skip to content

Category

machine learning

3,595 papers

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

WorldCupArena is presented, a dynamic benchmark for language models and deep-research agents that can be reused for future leagues and cups, and shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline.

Zhaokai Wang, T. Gui, Jiayuan Rao et al. · 1 citation · ⚡1
#machine learning Preprint Open access Sep 2026

The Value of Depth in Message Passing on Sparse Graphs: A Kesten-Stigum Dichotomy

How deep does a graph neural network need to be on a sparse graph? We study its purest statistical form: node classification on the sparse contextual stochastic block model (CSBM) with average degree $\Delta=O(1)$, whose local weak limit is a broadcast-labelled Poisson Galton-Watson tree. Prior work derived a message-passing classifier $h_\ell$ that aggregates from each vertex at distance $k\le\ell$ the attenuated evidence $2\operatorname{artanh}(\gamma^k t(X_v))$, with $\gamma$ the edge signal and $t$ a bounded likelihood-ratio transform of the feature. We prove that the value of depth is governed by a single number, the Kesten-Stigum ratio $\kappa=\gamma^2\Delta$. Below the threshold ($\kappa<1$), the error sequence is Cauchy at a geometric rate, $|\mathcal{E}(\ell)-\mathcal{E}(\ell')|\le C\kappa^{(\ell+1)/3}$ for all $\ell'>\ell$, so all layers beyond depth $O(\log(1/\epsilon))$ change the error by less than $\epsilon$; conversely, under mild regularity each sufficiently deep layer still flips the decision with probability at least $c\kappa^{\ell/2}$, the empirically sharp exponent. Above the threshold ($\kappa>1$), depth is geometrically productive: $\mathcal{E}(\ell)$ is driven to a branching-process floor of order at most $1/(\kappa-1)$ at any geometric rate $\kappa^{-s\ell}$, $s<1$ (this bound has content only for $\kappa>17$). No local classifier of any depth beats the universal floor $e^{-\Delta}\Phi(-\zeta)$ set by isolated roots ($\zeta$ the feature signal-to-noise ratio), while the first layer provably helps by an explicit total-variation amount. Simulations with an exact belief-propagation baseline on the same trees show that the pairwise rule's error curve is mildly non-monotone in $\ell$, so an optimal finite depth exists (an exact instance is certified in the appendix), while BP saturates strictly faster, at an effective per-layer ratio below $\kappa$ that we identify.

Aseem Raj Baranwal · 0 citations
#machine learning Preprint Open access Sep 2026

RegionFM: Interpretable Region-Based Brain MRI Classification Using Foundation Model Embeddings

Foundation models provide powerful representations for brain MRI analysis, but their predictions remain difficult to interpret in anatomically meaningful terms. Clinical assessment of brain MRI is commonly organized around anatomically defined structures and regional abnormalities, whereas conventional explanation methods typically produce voxel- or patch-level importance maps that do not explicitly quantify the contributions of individual brain regions. To address this mismatch, we propose RegionFM, an interpretable framework that integrates anatomical segmentation with brain MRI foundation-model embeddings. RegionFM first divides each MRI scan into anatomical regions and constructs a separate MRI volume for each region. A frozen foundation model then encodes each region into an embedding, and a region-additive logistic model combines these embeddings such that every anatomical region contributes an explicit scalar term to the final prediction. This formulation supports both subject-level and cohort-level analyses of regional contributions. We evaluate RegionFM on cognitive-impairment classification using embeddings from multiple pretrained brain MRI foundation models. The results show that RegionFM maintains performance comparable to less interpretable fine-tuning approaches while providing anatomically grounded explanations. Randomized embedding ablations yield near-chance performance, indicating that the predictions rely on meaningful structure captured by the foundation-model embeddings rather than simple feature statistics. Overall, RegionFM better aligns model explanations with anatomy-based clinical reasoning while maintaining competitive predictive performance.

Wei Zhang · 0 citations
#machine learning Preprint Open access Sep 2026

Posterior Variance Is a Constraint Map, Not an Error Map: Closed-Form Uncertainty for Radiative Gaussian Splatting in Sparse-View CT

Radiative Gaussian splatting reconstructs sparse-view CT fast and accurately, and recent work attaches per-Gaussian posteriors to yield per-voxel uncertainty maps. We ask what such a map actually measures: posterior variance is a data-constraint map, not an error map -- its alarms are trustworthy, its all-clears are not. Exploiting the strict linearity of X-ray rendering in the per-Gaussian densities, we derive a clamp-aware closed form that the unchanged rasterizer evaluates exactly in one forward pass, in volume and projection space: the infinite-sample limit of the sampling estimator of concurrent work, at ~8x lower cost. On the official 15-scene benchmark this uncertainty ranks true error on 14 of 15 scenes. Restricted to the object interior -- the tissue a clinician reads -- the ranking collapses (median Spearman 0.11, 0/15 pass), identically for a deep ensemble and for a strictly positive log-normal posterior: three constructions, two estimator families, no survivors. The mechanism is structural: about 90% of in-object error is bias that reproduces across retrainings, invisible to model disagreement; 73-81% of the full-volume correlation is carried by object/surround contrast; and an exactly solvable control puts the observed in-object ranking 4-5x below what a perfectly calibrated posterior with the same sigma-spread would score. The error scale, by contrast, is an engineering problem, and we solve it: reparameterizing the posterior contracts the cross-scene temperature spread from 19.3x to 2.6x, one scene-agnostic temperature transfers to unseen scenes (10/15 leave-one-scene-out), and the repaired scale tracks photon count at the Poisson-predicted -1/2 power. We distill evaluation practice that would have caught the illusion -- masked calibration, seed-wise bias decomposition, an exact-posterior reference -- and release all protocols, seeds and per-run evidence.

Chulin Zhao, Yiran Xu, Shu Liu · 0 citations

Runtime Safety Filtering for Learned Small UAS Separation Policies under GNSS Degradation

Experimental results show that action filtering provides negligible safety improvement, while observation filtering reduces near mid-air collisions by 90% and remains robust to the barrier function's tradeoff between separation distance and closing rate, and suggest that preserving the policy's decision authority outperforms overriding its actions with hand-designed constraints.

Alex Zongo, Peng Wei · 0 citations
#artificial intelligence Preprint Jul 2026

Cultural Bias Without a Cultural Self:A Disassociation Study of LLM's Persona and Bias

Language models prompted with cultural personas increasingly stand in for human respondents in cross-cultural research. Their responses separate personas cleanly, and that separation is read as evidence of a cultural point of view. We show that the separation is real, that the point of view is not, and that one criterion tells them apart. A trait is structure internal to one respondent that survives a change of measurement frame; a bias needs only group-specific item means. To test for the first, we represent a single response set as an Item--Dimension matrix and treat its correlation matrix as a point on the manifold of symmetric positive definite matrices. In humans this carries what a trait should: it reproduces across test--retest sessions sharing no items, order or context ($r=0.77$, $N=89$); on public NEO-PI-R data it identifies individuals at up to $76\%$ against a $0.4\%$ chance level ($N=263$); and it predicts GPA ($R^2=0.281$, $p=0.003$) where BigFive aggregates from the same responses predict nothing ($R^2=0.018$). In four frontier LLMs it returns nothing. Persona structure is readable only while every instance shares one item order: give each its own order and separation falls from $94.7\%$ to chance, while realigning instances to \emph{any} shared random order restores it to $82$--$84\%$. Responses generated independently item by item, with no latent structure, reproduce the entire pattern. The cultural signal is a group template, not a property of any instance, and alignment regimes differ only in which stereotype survives on the surface.

Yu Yuan · 0 citations

Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery

Self-Organized Conformal Prediction (SOCP), a calibration scheme that discovers input-space groups with an unsupervised Self-Organizing Map (SOM) trained without calibration labels, is introduced, providing a concise route to group-local calibration without supervised partitions or predictor retraining with a diagnostic toolkit.

Louis Berthier, A. Shokry, M. Moreaud et al. · 0 citations

MultiHashFormer: Hash-based Generative Language Models

This paper proposes MultiHashFormer, a new framework that allows hash-based autoregression that consistently outperforms standard Transformer LMs across multiple benchmarks and shows that the model handles multilingual vocabulary expansion with a constant parameter footprint without any modifications.

Hui Xue, Atsuki Yamaguchi, Nikolaos Aletras · 0 citations
#machine learning Preprint Open access Sep 2026

FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel Dynamics

Event cameras asynchronously report brightness changes with microsecond-level temporal resolution, but real event data remain difficult to collect at scale because specialized sensors, careful synchronization, and task-specific annotations are required. Event-camera simulation is therefore important to event-based vision tasks. Most practical simulators build on contrast-threshold event generation, some with additional filtering, stochastic noise, or hand-tuned sensor parameters. While effective, such formulations often simplify the temporal structure produced by the lifecycle of each pixel, which can distort event timing and weaken downstream transfer. We introduce FracEvent, an event simulator that models this pixel-level lifecycle with fractional-relaxation voltage dynamics. Given a log-intensity trajectory, FracEvent drives a compact stack of relaxation modes, combines their responses into a voltage state, emits ON/OFF events by localizing threshold crossings on the continuous voltage trajectory, and updates the reference while retaining the underlying memory modes. This retained state links residual voltage response to later event timing. We evaluate FracEvent through event-stream comparison and downstream transfer on image reconstruction and optical flow estimation. Across multiple datasets, FracEvent improves the temporal structure of generated events and achieves stronger downstream-transfer results than competing simulator baselines, showing its practical value for event-camera simulation.

Langyi Chen, Chuanzhi Xu, Haoxian Zhou et al. · 0 citations

In LLM Reasoning, there is Irrationality on top of Value Misalignment

It is argued that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning, and the utility discrepancy between a model's deployed reasoning strategy and its rational counterpart whose responses maximise utility in the steepest direction is formalised.

Kejiang Qian, Feng-Xiang He · 0 citations

PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

Comprehensive empirical evaluations demonstrate PEAR significantly improves average accuracy over the strongest debate baselines, and theoretically characterize PEAR as an equivariant sparse router: it preserves accuracy under agent relabeling while reducing routing complexity and improving generalization.

Yang Feng, Ziwei Xu, Xia Hu et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.