Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a...
Shuai-Jun Liu, Cheng-Ju Wu, Qi-Fu Wen et al.· 0 citations
Vision-language-action (VLA) models have achieved strong performance in embodied manipulation, but still lack a clear mechanism to balance behavioral stability with task-semantic sensitivity. We identify two complementary failure modes. Under task-preserving changes, where task semantics remain unchanged but scene appe...
Shuai-Jun Liu, Fei-Yang You, Cheng-Ju Wu et al.· 0 citations
Embodied agents replan frequently to recover from execution drift, partial observability, and coordination hazards, but each LLM-based replanning call can consume an accumulated textual context that grows over time and across agents. Once this context becomes large, replanning latency develops heavy tails and can miss...
Shuaijun Liu, Feiyang You, Xingwei Chen et al.· 0 citations
It is proved that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM, and isolates two missing links between internal update gaps and predictive cost.
Qi-Fu Wen, Shuai Liu, Zihan Zhou et al.· 0 citations
Token skipping is a widely used training-free way to accelerate vision--language--action (VLA) models by bypassing computation for most visual tokens at each control step according to a gate. When the next gate is harvested from the previous accelerated forward, however, the tokens skipped at one step are also the ones...
Qi Luo, Shuaijun Liu, Hao Zhao et al.· 0 citations
This work introduces PortBench, a benchmark spanning six heterogeneous asset classes from 2015 to 2025, and introduces two metrics: a dual-layer correlation score for inter-class hedging and intra-class concentration, and CEPS, which quantifies how reasoning errors compound across pipeline stages.