Personalized large language models are often expected to follow explicit style instructions, yet we find that such instructions can undermine the user-specific characteristics that personalization methods aim to preserve. We call this failure mode personalization collapse: explicit style control can conflict with impli...
Yutong Song, Jiang Wu, Shaofan Yuan et al.· 0 citations
Adapting large language models to individual users remains challenging due to the tension between fine-grained personalization and scalable deployment. We present CARD, a hierarchical framework that achieves effective personalization through progressive refinement. CARD first clusters users according to shared stylisti...
Yutong Song, Jiang Wu, Weijia Zhang et al.· 0 citations
Chemical reasoning inherently integrates visual, textual, and symbolic modalities, yet existing benchmarks rarely capture this complexity, often relying on simple image-text pairs with limited chemical semantics. As a result, the actual ability of Multimodal Large Language Models (MLLMs) to process and integrate chemic...
Zhiyuan Huang, Baichuan Yang, Zikun He et al.· 0 citations
Experiments show that the NuSA-CL framework not only effectively preserves zero-shot transfer capabilities but also achieves highly competitive performance on continual learning benchmarks, positioning NuSA-CL as a practical and scalable solution for continually evolving zero-shot VLMs in real-world applications.
Yujin Jo, Taesup Kim· arXiv.org· 2 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Retrieval-Augmented Generation (RAG) systems empower large language models (LLMs) with external knowledge, yet struggle with efficiency-accuracy trade-offs when scaling to large knowledge graphs. Existing approaches often rely on monolithic graph retrieval, incurring unnecessary latency for simple queries and fragmente...
Ruiyi Yang, Hao Xue, Imran Razzak et al.· 0 citations
Artificial intelligence is rapidly advancing scientific discovery, but this progress carries risks of misuse, such as the creation of harmful substances, or circumvention of established regulations. In this paper, we first demonstrate the risks by highlighting real-world examples of AI misuse in chemical science, which...
Jiyan He, Haoxiang Guan, Weitao Feng et al.· 0 citations
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, o...
Yi-Ran Wang, Xingyilang Yin, Junfu Pu et al.· 0 citations
WorldCrafter is a video world model that learns a camera-queryable implicit 3D-aware memory that enables streaming scene exploration from a single input image or text prompt and shows substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploratio...
Wang-Bo Yu, Kunhao Liu, Wen-Bo Hu et al.· 1 citation
Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a...
Hao-Ran Yuan, Ze-Kai Wang, Boning Shao et al.· 0 citations
Dormant tree pruning is labor-intensive yet essential for maintaining modern high-productivity fruit orchards. In this work, we focus on pruning of modern planar tree training systems - V-Trellis apples and UFO cherries - where trunks and primary branches are trained into approximately planar walls. We introduce an end...
Abhinav Jain, Cindy Grimm, Stefan Lee· 0 citations
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accuracy and compute. Potential fields learned directly from images remove that dependency but inherit the classical weakness...
Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase et al.· 0 citations
LLM inference on edge devices is constrained by computational and memory resources, making efficient autoregressive decoding challenging. Speculative decoding alleviates this bottleneck by generating tokens with a smaller draft model and verifying multiple tokens in parallel with a batched target model pass. However, v...
Gabriele Tombesi, William Baisi, Je Yang et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.