Aug 2026· ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP)· 0 citations
TL;DR
Five contributions addressing key challenges in LAM research are brought together, including safety and robustness against jailbreak and adversarial attacks, semantic-perceptual integration for robotic manipulation, efficient deployment on edge devices, and natural-language-driven decision-making for networked systems.
Abstract
Large Action Models (LAMs) extend the capabilities of AI systems beyond text generation toward perception, reasoning, and action, enabling applications across robotics, autonomous systems, smart manufacturing, healthcare, and the Internet of Things. This Special Issue of ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) brings together five contributions addressing key challenges in LAM research, including safety and robustness against jailbreak and adversarial attacks, semantic-perceptual integration for robotic manipulation, efficient deployment on edge devices, and natural-language-driven decision-making for networked systems. Together, these papers span the theoretical, implementation, and application dimensions of LAMs, offering both practical solutions and a foundation for future research toward LAM-based systems that are safe, efficient, and reliably grounded in action.
This framework surveys manipulation, navigation, locomotion, autonomous driving, and general embodied learning, tracing technical progressions, clarifying capability requirements, and examining datasets, benchmarks, and evaluation protocols.
Nan-Jie Yao, Hao Wang, Chong Cheng et al.· 0 citations
Chain-of-Thought (CoT) reasoning is increasingly incorporated into Vision-Language-Action (VLA) models, yet it can degrade the performance of stronger embodied agents. We investigate this capability-dependent effect by distinguishing explicit reasoning from latent decision computation, i.e., perception-grounded computa...
Yuan-Peng Lin, Zi-Yu Zhou, Jin-Long Zhao et al.· 0 citations
Robots operating in open environments act under partial observability, physical constraints, and dynamic task contexts. Beyond mapping observations and language instructions to actions, they must anticipate how candidate actions may affect future states and task-relevant outcomes. Recent advances in world models, video...
Zu-Xing Lu, Hong-Jia Zhai, Guan-Zhi Wang et al.· 1 citation
World Action Models (WAMs) have emerged as a promising paradigm for generalizable robot control. Despite the growing number of WAM systems, existing works often introduce unified systems that bundle together multiple design choices, such as architecture and training strategy, making it difficult to isolate individual c...
Chao Tang, Haoqing Wang, Zi-Lang Cen et al.· 0 citations
General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize...
Hao-Jian Huang, Ze-Xi Li, Ju-Hao Guo et al.· 1 citation
Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce large language models (LLMs), vision-language models (VLMs), or language-derived semantics at different stages, but reported gains are difficult to interpret when task formula...