Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a la...
Ming-Qing Yuan, Xiao-Bo Liang, Jun-Wei Yang et al.· 0 citations
Flux Attention is introduced, a context-aware framework that dynamically optimizes attention computation at the layer level by integrating a lightweight Layer Router into frozen pretrained LLMs, which adaptively routes each layer to FA or SA based on the input context.
Quantong Qiu, Zhiyi Hong, Yi Yang et al.· arXiv.org· 0 citations
D-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online to adapt the domain mixture online.
Zechen Sun, Zhi-Wei Zhang, Fei Zhao et al.· 1 citation
The ProgramTab framework is proposed, which guides LLMs employing in-context learning to perform tabular data preprocessing with Python code, as well as the momentous contents extraction with row and column extraction and SQL generation, demonstrating that the ProgramTab framework effectively deals with table-based rea...
Pei Guo, Enjie Liu, Yunzhi Tan et al.· arXiv.org· 0 citations
This work introduces MMLongEmbed, the first comprehensive benchmark for evaluating MEMs in long-context scenarios, and finds that current architectures rely heavily on superficial feature matching and struggle to capture deep semantic and structural dependencies.
This work proposes SABER, a training-free framework for stability-aware early exit via adversarial branch probing, and shows that SABER reduces reasoning token consumption by 30.2% on average while maintaining competitive accuracy with full-length reasoning.
Wanzhe Cheng, Hai-Yang Xiang, Jun-Tao Li et al.· 1 citation
This work provides a scalable and effective framework for extending RLVR beyond the limitations of pattern-based verification to complex, noisy, real-world domains, and generalizes strongly to seven out-of-distribution benchmarks.
Yi Su, Dian Yu, Linfeng Song et al.· Annual Meeting of the Associ...· 1 citation
A Hierarchical Online Memory Exploration and Reasoning framework that mirrors the multi-scale structure of long videos, and consistently lifts three various LLM backbones, indicating a model-agnostic structural capability for grounded retrieval over long videos.
As the context window of Large Language Models (LLMs) continues to expand, the data required to effectively train and evaluate these capabilities remains underexplored. With existing research primarily focuses on architectural optimization, there is a need for a systematic, data-centric review. This survey bridges th...
Zechen Sun, Yu-Yang Sun, Zhao-yu Su et al.· Transactions of the Associat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.