Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored. Existing benchmarks rely on action-specific annotation schemes and focus primarily on final assessment outputs, offering limited in...
Kaili Zheng, Kaiwen Wang, Xun Zhu et al.· 0 citations
A controlled relay testbed is introduced in which briefs of twelve programmatic atomic facts are re-encoded hop by hop in five formats over six hops, scored against programmatic ground truth by a fixed strong grader, across two relay-capability tiers, a cognitive-load condition, and a paired-fork error injection.
Bounded LLM reasoning can provide a practical metacognitive control layer over ongoing learning processes and is suggested that bounded LLM reasoning can provide a practical metacognitive control layer over ongoing learning processes.
Financial decision-makers face more information than they can directly inspect, making context compression necessary. Yet when large language models (LLMs) compress financial source material, they can alter the investment judgment supported by the original source. We frame this problem as information fidelity: compress...
Hoyoung Lee, Suhwan Park, Seunghan Lee et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Adaptive hierarchical systems accumulate routing state as they learn which components to select. We show that this state already defines a coherent attribution over the hierarchy. A leaf receives the product of the local routing weights on its path, while an internal node receives the corresponding prefix product. The...
Large Language Models (LLMs) are increasingly deployed as scientific AI as- sistants, and a growing body of benchmarks evaluates their capabilities across knowledge retrieval, reasoning, code generation, and tool use. These evaluations, however, typically assume the scientific problem is already well-posed, whereas pra...
Nithin Somasekharan, Youssef Hassan, Shiyao Lin et al.· 0 citations
Ambiguity is an inherent property of natural-language agent specifications. When a system prompt leaves behaviour underdetermined, identical inputs follow divergent execution paths and produce inconsistent outcomes. The standard remedy is prompt optimisation: propose candidate prompts, run the agent to score them, and...
Yuval David, Fabiana Fournier, Lior Limonad et al.· 0 citations
In this work, we propose Oph-Guid-RAG, a multimodal visual RAG system for ophthalmology clinical question answering and decision support. We treat each guideline page as an independent evidence unit and directly retrieve page images, preserving tables, flowcharts, and layout information. We further design a controllabl...
Ordinary decentralized multi-agent reinforcement learning presents each focal agent with a continual learning problem: peer updates change its induced rewards and dynamics even when the joint Markov game is stationary. We connect the lifetime of success-conditioned reusable structure to peer learning and policy reuse....
Large language models (LLMs) display a unified "general factor" of capability across 10 benchmarks (a finding confirmed by our factor analysis of 156 models), yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cog...
Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh· 0 citations
An architectural taxonomy is introduced that decomposes multi-agent LLM frameworks along five dimensions: orchestration, memory, planning interfaces, specialization, and communication topology, and a controlled empirical study is conducted, fixing the underlying LLM and varying only architectural design choices.
Abdelghny Orogat, Ana Rostam, Essam Mansour· 4 citations
While open sourced Vision-Language Models (VLMs) have proliferated, selecting the optimal pretrained model for a specific downstream task remains challenging. Exhaustive evaluation is often infeasible due to computational constraints and data limitations in few shot scenarios. Existing selection methods fail to fully a...
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.