VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjective qualities with a structured 5-stage training curriculum.
A novel framework that aligns multi-trajectory supervision with policy optimization, and introduces two complementary mechanisms: feasibility-first advantage assignment and dynamic distillation to ensure that expanded trajectory supervision is effectively absorbed during policy optimization.
Tian Zhang, Zhuo Huang, Hong-Rui Ye et al.· 0 citations
This work presents an open-source transformer implementation for uncropped full-key attacks which uses the standard transformer encoder backbone, adapting only the input and output layers to the side-channel setting.
This work renders each administrative unit as a single polygon-masked satellite image and treats tract-level population estimation as a sequence-modeling problem over its image patches, pairing each tract image directly with its population label and eliminating the disaggregation step entirely.
Jackson R. Ye, Alexandr V. Morozov· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work study how a fitted classifier and an LLM can be combined for credit-default prediction, and recommends a simple classifier-guided prompt for LLM-based credit prediction.
A statistical framework connecting token prediction with representation geometry, encoder approximation, and downstream performance is developed, introducing a self-consistency principle showing that repeated applications of a shared representation block can progressively refine the contextual representation without introducing additional block parameters.
This work theoretically shows that directly leveraging PRM score is vulnerable to verifier noise through an extreme-value effect: non-viable prefixes become more likely to receive spuriously high scores as reasoning depth increase, leading to a training-free robust process supervision method that preserves promising alternatives when step-level scores are noisy.
Balance of Benchmarks (BoB) is introduced, which embeds benchmark descriptions and assigns each benchmark an inverse-density semantic weight, providing a principled foundation for task-aware and multiplicity-robust model evaluation.
A unified probabilistic framework that jointly addresses missing data, measurement error, and population heterogeneity utilizing deep latent variable representation is proposed that integrates a novel hierarchical tree-routed variational autoencoder with pattern-aware latent representations and calibration-based denoising.
Yasin Khadem Charvadeh, Grace Y. Yi, Mithat Gönen et al.· 0 citations
A multi-solver disagreement reward using a heterogeneous ensemble varying in model capacity and sampling temperature is proposed, which enables the Challenger to discover questions targeting true capability boundaries, producing a curriculum that forces downstream Solvers to develop robust reasoning strategies generalizing across problem types.
This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the evaluation objectives.
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al.· 0 citations
The results show that PURGE consistently reduces hallucinations and spurious-correlation-driven errors while maintaining or improving overall performance in most evaluated settings, providing both a reusable evaluation protocol and an effective mitigation framework for more reliable LVLMs.
Aditi Sarker, Nazreen Shah, Rafi Ibn Sultan et al.· 0 citations
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.