Skip to content

Category

machine learning

4,920 papers

#artificial intelligence Preprint Aug 2026

VIBE: Video Instruction-aligned Background music gEneration

VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjective qualities with a structured 5-stage training curriculum.

Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving

A novel framework that aligns multi-trajectory supervision with policy optimization, and introduces two complementary mechanisms: feasibility-first advantage assignment and dynamic distillation to ensure that expanded trajectory supervision is effectively absorbed during policy optimization.

Tian Zhang, Zhuo Huang, Hong-Rui Ye et al. · 0 citations
#machine learning Preprint Aug 2026

A Hybrid State-Space Approach for Census-Tract Population Estimation

This work renders each administrative unit as a single polygon-masked satellite image and treats tract-level population estimation as a sequence-modeling problem over its image patches, pairing each tract image directly with its population label and eliminating the disaggregation step entirely.

Jackson R. Ye, Alexandr V. Morozov · 0 citations
#machine learning Preprint Aug 2026

Learning Representations through Token Prediction: Geometry, Approximation, and Downstream Guarantees

A statistical framework connecting token prediction with representation geometry, encoder approximation, and downstream performance is developed, introducing a self-consistency principle showing that repeated applications of a shared representation block can progressively refine the contextual representation without introducing additional block parameters.

Shu-Lei Wang · 0 citations
#artificial intelligence Preprint Aug 2026

Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the Guide

This work theoretically shows that directly leveraging PRM score is vulnerable to verifier noise through an extreme-value effect: non-viable prefixes become more likely to receive spuriously high scores as reasoning depth increase, leading to a training-free robust process supervision method that preserves promising alternatives when step-level scores are noisy.

Taejong Joo, Diego Klabjan · 0 citations
#machine learning Preprint Aug 2026

A Deep Latent Variable Framework for Jointly Modeling Missingness, Measurement Error, and Heterogeneity

A unified probabilistic framework that jointly addresses missing data, measurement error, and population heterogeneity utilizing deep latent variable representation is proposed that integrates a novel hierarchical tree-routed variational autoencoder with pattern-aware latent representations and calibration-based denoising.

Yasin Khadem Charvadeh, Grace Y. Yi, Mithat Gönen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula

A multi-solver disagreement reward using a heterogeneous ensemble varying in model capacity and sampling temperature is proposed, which enables the Challenger to discover questions targeting true capability boundaries, producing a curriculum that forces downstream Solvers to develop robust reasoning strategies generalizing across problem types.

Vinoth Selvendran, Zhan-Ming Zhang · 0 citations
#artificial intelligence Preprint Aug 2026

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the evaluation objectives.

Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language Models

The results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.

Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan et al. · 0 citations

From tech blogs

See all →
GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.