This paper articulates an alternative and complementary paradigm: perception and synthesis as distinct conditionings within a shared Generalist Open-World Temporal Perception Architecture (GOWTPA), positioning generalist temporal perception as a potential foundation layer for broader multimodal intelligence and physical AI.
Abstract
The next generation of artificial intelligence systems will likely be natively temporal and multimodal in both inputs and outputs: able to converse, perceive, predict, reason, and synthesize through a shared world representation. Realizing this requires a temporal perceptual substrate integrating sensory streams, language, and structured outputs within a multimodal world model. We seek a generalist open-world perceptual system that represents biological forms, natural physical structures, and artifacts, and their interactions, as a coherent, temporally persistent process. The model should infer geometry, articulation, semantics, interaction structure, and uncertainty from raw multimodal streams; maintain identity through occlusion and viewpoint change; generalize across species, forms, mechanisms, and materials; and abstain or expand its ontology when encountering the unknown. The objective is a structured world state supporting understanding, prediction, counterfactual reasoning, and controllable synthesis. Recent work suggests that some cross-modal and reasoning-like capabilities can emerge from large-scale generative video pretraining, reminiscent of language-model scaling. Yet these capabilities are often accessed through language probes or expressed through photorealistic video, leaving explicit semantic, geometric, or temporal structure largely unexposed. This paper articulates an alternative and complementary paradigm: perception and synthesis as distinct conditionings within a shared Generalist Open-World Temporal Perception Architecture (GOWTPA). Recognition, structured prediction, and simulation arise as different conditionings of the same generative substrate, while reasoning and embodiment-specific policies build upon the resulting world state. This positions generalist temporal perception as a potential foundation layer for broader multimodal intelligence and physical AI.
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world models has significantly enhanced the capabilities of robot learning systems, enab...
Shaunak A. Mehta, Ananya Hazarika, Haochen Zhang et al.· Trans. Mach. Learn. Res.· 0 citations
A causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes that learns unified latent states that integrate vision, language, and action, and strengthens generalization across diverse navigation contexts.
Xuan Yao, Junyu Gao, Chang-Sheng Xu· IEEE Transactions on Pattern...· 0 citations
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...
Wen-Xuan Song, Jia-Yi Chen, Jing-Bo Wang et al.· 1 citation
Embodied Semantic Grounding (ESG) is proposed, a framework that equips LLMs with consequence-aware text representations that extends affordance-grounding with consequence-level representations of post-event environmental functionality.
Manaswi Kulahara, Khadija Parwez, Faisal Alhwikem et al.· Computers, Materials & C...· 0 citations
This Review synthesizes recent progress in Embodied AI and articulate Intent-Driven Embodied Artificial Intelligence (IDEAI) as a system-level organizing framework in which intent functions as an explicit, revisable, and verifiable mediating construct between human goals, environmental constraints, and agent behavior.
Nanning Zheng· National Science Review· 0 citations
A scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing is introduced.