Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive phenomenon exists: post-training models for structured reasoning trajectories results in marginal or even negative gains compared to post-tra...
Ruixiang Mao, Xiangnan Ma, Dan Chen et al.· 0 citations
Across five real-robot settings with all policies running on the same Jetson AGX Thor, HybridFlow improves normalized task performance over 16-step Diffusion Policy by 13-68 points with approximately eightfold lower action-generation latency.
Zhen Dong, Fu-Lin Chen, Jin-Na Fu et al.· 0 citations
Autoregressive speech synthesis often adopts a left-to-right order, yet generation order is a modelling choice. We investigate decoding order through masked diffusion framework, which progressively unmasks positions and allows arbitrary decoding orders during training and inference. By interpolating between identity an...
Human identity-preserving text-to-video generation remains challenging under large changes in viewpoint, facial expression, illumination, and motion. Existing methods condition the generator on a single reference portrait, but a static image cannot capture how identity-bearing cues evolve across views and expressions,...
Yixuan Lai, He Wang, Kun Zhou et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large language models (LLMs) and vision language models (VLMs), opens the possibility of teams that infer mission-relevant semantics and subtasks given high-level natural language spec...
Zachary Ravichandran, Fernando Cladera, Ankit Prabhu et al.· 0 citations
Agentic artificial intelligence is a candidate enabler of Level-4 autonomy in sixth-generation (6G) networks, but agents reasoning over a shared memory inherit its distortions. We study cross-domain radio access network (RAN)--edge orchestration in which a RAN agent minimizing energy and an edge agent minimizing latenc...
Hatim Chergui, Farhad Rezazadeh, Miguel Catalan Cid et al.· 0 citations
Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to date focus on gene-level information. We present one of the first systematic frameworks to generate genetic variant-level embeddings across the entire human genome. Using curated...
Hongqian Niu, Jordan Bryan, Jacob Williams et al.· 0 citations
A novel instrumental variable estimator is developed that accommodates multivariate outcomes, sparse networks, and multidimensional latent homophily and is shown to be $\sqrt{N}$-consistent and asymptotically normal under sparsity conditions that relax dense-network assumptions prevalent in the peer effect literature.
Shanjukta Nath, Jiwon Hong, Jae Ho Chang et al.· 0 citations
Cross-chain bridges enable asset and state transfers across heterogeneous blockchains, but their complex cross-domain interactions introduce new attack surfaces that are difficult to monitor using traditional single-chain analysis methods. Existing approaches often focus on isolated on-chain behaviors and fail to captu...
Dan Lin, Shunfeng Lu, Ziyan Liu et al.· 0 citations
Large Language Models (LLMs) have been widely adopted in commercial code completion engines, significantly enhancing coding efficiency and productivity. However, even functionally correct LLM-generated code may exhibit non-functional quality issues that violate coding standards and best practices, such as poor style an...
Liang Lu, Yuan Jiang, Christoph Treude· 0 citations
Addressing the challenge of ensuring safety in ever-changing and unpredictable environments, particularly in the swiftly advancing realm of autonomous driving in today's 5G wireless communication world, we present Navigation Secure (NavSecure). This vision-based navigation framework merges the strengths of world models...
Hong Ding, Ziming Wang, Yi Ding et al.· 0 citations
In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the...
Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor et al.· 0 citations
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.