Skip to content

Category

machine learning

3,595 papers

#machine learning Preprint Open access Sep 2026

PhyloGFN: Phylogenetic inference with generative flow networks

Phylogenetics is a branch of computational biology that studies the evolutionary relationships among biological entities. Its long history and numerous applications notwithstanding, inference of phylogenetic trees from sequence data remains challenging: the high complexity of tree space poses a significant obstacle for the current combinatorial and probabilistic techniques. In this paper, we adopt the framework of generative flow networks (GFlowNets) to tackle two core problems in phylogenetics: parsimony-based and Bayesian phylogenetic inference. Because GFlowNets are well-suited for sampling complex combinatorial structures, they are a natural choice for exploring and sampling from the multimodal posterior distribution over tree topologies and evolutionary distances. We demonstrate that our amortized posterior sampler, PhyloGFN, produces diverse and high-quality evolutionary hypotheses on real benchmark datasets. PhyloGFN is competitive with prior works in marginal likelihood estimation and achieves a closer fit to the target distribution than state-of-the-art variational inference methods. Our code is available at https://github.com/zmy1116/phylogfn.

Mingyang Zhou, Zichao Yan, Elliot Layne et al. · 0 citations
#machine learning Preprint Jun 2023

Deep graph kernel point processes over networks

A new point process model for discrete-event data over networks is presented, built upon Hawkes' classic influence-kernel formulation to capture the effects of historical events on the occurrence of future events.

Zheng Dong, Matthew P. Repasky, Xiuyuan Cheng et al. · 5 citations
#machine learning Preprint Aug 2026

Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting

Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality.

Xinwei Qiang, Xiang Fang, Changli Chen et al. · 0 citations
#machine learning Preprint Aug 2026

QuantumBoostNet: Hybrid Classical-Quantum Cardiac View Identification

The hybrid classical-quantum architecture QuantumBoostNet is proposed, which combines a classical backbone with two heads: one classical and one quantum, a parametrized 10-qubit quantum circuit, which outperforms the implemented baselines under matched training conditions.

Mihai Udrescu-Milosav, S. Jura, M. Udrescu et al. · 0 citations
#machine learning Preprint Aug 2026

A Unified Framework for Fair and Personalized Decentralized Learning under Communication Constraints

A new algorithm DMFL-SQ is proposed, a decentralized multi-task learning algorithm that couples personalized model training over a communication graph with an agnostic mixture fairness objective, while reducing communication through sparsification, quantization, and event-triggered synchronization.

Krishnendu S. Tharakan, Carlo Fischione · 0 citations
#artificial intelligence Preprint Aug 2026

ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations

CON decomposition is introduced, which quantifies how much of a layer's variance each concept explains given all other concepts and the outcome, and how much none of them explains, yielding layer-comparable, calibrated scores that suppress false positives.

R. Rane, Marco Simnacher, Manuel Pfeuffer et al. · 0 citations
#artificial intelligence Preprint Aug 2026

It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning

This paper argues that the current main approaches in multi-objective RL (SER and ESR), and successor features, are insufficient, and motivates that this can indeed be the case by an example, leading to a new perspective, and a significant and non-trivial gap in the literature.

L. P. J. Mertens, L. N. Alegre, Florent Delgrange et al. · 0 citations
#machine learning Preprint Aug 2026

Frequency-aware forecasting for short-term typhoon gust prediction

WDANet, a frequency-aware forecasting framework that integrates stationary wavelet decomposition, a Feature-wise Linear Modulation strategy, and a dual-branch encoder-decoder architecture, enabling separate modeling of trend and fluctuation components is proposed, highlighting its potential for offshore wind power operation, disaster warning, and risk mitigation.

Xue-Fei Wang, Tingting Liu, Heng Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment

We consider the problem of learning a mixture of $k$ Plackett-Luce models given multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, uncovering mixture models is theoretically unidentifiable when $k$ exceeds $m/2$, where $m$ is the length of a ranking. We propose an efficient implementation to address this limitation, which involves first augmenting the rankings to a larger size by generating new responses from a base language model, followed by a gradient-based estimation to reduce inference cost in the input embedding space. Based on this procedure, we then design an expectation-maximization algorithm with these two steps to fit a mixture of Plackett-Luce models, called MoPLEx. Extensive experiments are conducted to verify this approach. First, we show that the gradient-based approximation estimates true probabilities with less than 5% error on models with up to 34 billion parameters. Second, we show that MoPLEx improves clustering and ranking accuracy by an average of 43.7% and 15.2% over baselines using single ranking and mixtures of Bradley-Terry models, on preference optimization datasets. These results demonstrate the effectiveness of MoPLEx for tackling multi-way rankings from heterogeneous preferences through measuring alignment between gradients.

Dongyue Li, Ziniu Zhang, Lu Wang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning

DeMMO is proposed, an interpretable framework for longitudinal, multi-disease, and multi-outcome learning that infers signed relations directly from learned longitudinal DMO-outcome mappings, thereby enabling selective information sharing across cohorts without requiring paired participants.

Menghui Zhou, Zhipeng Yuan, V. Lanfranchi et al. · 0 citations
#artificial intelligence Preprint Aug 2026

FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

A novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs and a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric.

G. Lee, Ji Liu, Juncheng Jia et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.