Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks...
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $\pi_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces subst...
Minjae Oh, Yoonah Park, Jongwon Lim et al.· 0 citations
The results suggest that the representation space in which diffusion operates should itself be treated as a first-class design variable for continuous diffusion in language models.
Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear when these choices...
Large language model (LLM) agents assist users with everyday tasks that can be completed in many reasonable ways. Even when their answers are useful, how agents carry out these tasks may not match users'preferences and needs. For example, agents differ in whether they ask clarifying questions or search the web. We intr...
Woojung Song, Hoyeol Yang, Jeonghoon Shim et al.· 0 citations
Personalization requires language models to capture user-specific knowledge from a growing user history. Existing context-based approaches incur increasing inference costs as user history accumulates and rely on separate retrieval or summarization stages, while parametric-based approaches often require reconstructing u...
Sejun Park, Hyo-Eun C. Bhang, Hyein Jeong et al.· 0 citations
LLM agents use tools to access information and perform computations beyond their parametric knowledge. Existing tool-use benchmarks evaluate whether agents select and call the right tools, assuming that tool returns are reliable. However, tools can return plausible but incorrect outputs. We evaluate 14 models with thre...
Hoyeol Yang, Woojung Song, Taewon Kim et al.· 0 citations
SpokenUS is presented, a spoken user simulator grounded in TOD that decides when to speak through a dedicated turn-taking head that achieves comparable goal coverage to much larger models while substantially outperforming all baselines in human MOS.
Jonggeun Lee, Junseong Pyo, Jeongmin Park et al.· arXiv.org· 0 citations
A conversational diagnosis system that explores a diagnostic knowledge graph to reason in two steps, generating diagnostic hypotheses from the dialogue context and verifying hypotheses through clarifying questions, which are repeated until a final diagnosis is reached.
Large language models (LLMs) achieve strong performance on mathematical reasoning benchmarks, yet the mathematically meaningful skills underlying their reasoning remain underexplored. We introduce \texttt{SHAPE}, a framework that analyzes Chain-of-Thought (CoT) trajectories through two lenses developed in mathematics e...
Jonghyun Song, Sangjun Song, Minjae Oh et al.· 0 citations
It is suggested that diffusion post-training selectively preserves or reorganizes inherited computation according to task structure, rather than uniformly replacing autoregressive mechanisms.
MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.
Joonki Min, Chaeyun Kim, H. Choi et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.