A new point process model for discrete-event data over networks is presented, built upon Hawkes' classic influence-kernel formulation to capture the effects of historical events on the occurrence of future events.
Zheng Dong, Matthew P. Repasky, Xiuyuan Cheng et al.· 5 citations
Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality.
Xinwei Qiang, Xiang Fang, Changli Chen et al.· 0 citations
The hybrid classical-quantum architecture QuantumBoostNet is proposed, which combines a classical backbone with two heads: one classical and one quantum, a parametrized 10-qubit quantum circuit, which outperforms the implemented baselines under matched training conditions.
Mihai Udrescu-Milosav, S. Jura, M. Udrescu et al.· 0 citations
A new algorithm DMFL-SQ is proposed, a decentralized multi-task learning algorithm that couples personalized model training over a communication graph with an agnostic mixture fairness objective, while reducing communication through sparsification, quantization, and event-triggered synchronization.
Krishnendu S. Tharakan, Carlo Fischione· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
CON decomposition is introduced, which quantifies how much of a layer's variance each concept explains given all other concepts and the outcome, and how much none of them explains, yielding layer-comparable, calibrated scores that suppress false positives.
R. Rane, Marco Simnacher, Manuel Pfeuffer et al.· 0 citations
This paper argues that the current main approaches in multi-objective RL (SER and ESR), and successor features, are insufficient, and motivates that this can indeed be the case by an example, leading to a new perspective, and a significant and non-trivial gap in the literature.
L. P. J. Mertens, L. N. Alegre, Florent Delgrange et al.· 0 citations
WDANet, a frequency-aware forecasting framework that integrates stationary wavelet decomposition, a Feature-wise Linear Modulation strategy, and a dual-branch encoder-decoder architecture, enabling separate modeling of trend and fluctuation components is proposed, highlighting its potential for offshore wind power operation, disaster warning, and risk mitigation.
Xue-Fei Wang, Tingting Liu, Heng Zhang et al.· 0 citations
We consider the problem of learning a mixture of $k$ Plackett-Luce models given multi-way ranking responses from annotators that may represent heterogeneous underlying preferences. This problem has many applications in AI alignment and preference optimization. Prior work has studied mixtures of Bradley-Terry models from pairwise comparisons. However, uncovering mixture models is theoretically unidentifiable when $k$ exceeds $m/2$, where $m$ is the length of a ranking. We propose an efficient implementation to address this limitation, which involves first augmenting the rankings to a larger size by generating new responses from a base language model, followed by a gradient-based estimation to reduce inference cost in the input embedding space. Based on this procedure, we then design an expectation-maximization algorithm with these two steps to fit a mixture of Plackett-Luce models, called MoPLEx. Extensive experiments are conducted to verify this approach. First, we show that the gradient-based approximation estimates true probabilities with less than 5% error on models with up to 34 billion parameters. Second, we show that MoPLEx improves clustering and ranking accuracy by an average of 43.7% and 15.2% over baselines using single ranking and mixtures of Bradley-Terry models, on preference optimization datasets. These results demonstrate the effectiveness of MoPLEx for tackling multi-way rankings from heterogeneous preferences through measuring alignment between gradients.
Dongyue Li, Ziniu Zhang, Lu Wang et al.· 0 citations
DeMMO is proposed, an interpretable framework for longitudinal, multi-disease, and multi-outcome learning that infers signed relations directly from learned longitudinal DMO-outcome mappings, thereby enabling selective information sharing across cohorts without requiring paired participants.
Menghui Zhou, Zhipeng Yuan, V. Lanfranchi et al.· 0 citations
A novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs and a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric.
A depth hierarchy for ReLU neural networks in which every additional ReLU layer can save exponentially many neurons is proved, and the first exponential separation for ReLU networks between two fixed depths whose shallower network has depth at least $3 is proved.
CatchBench puts one auditor's question to three information states: the declared configuration before a run (PRE), a growing prefix of its trace (LIVE), and the finished trace (POST), which none scores all three under one task-method interface.
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.