Skip to content
Preprint

Generalist Open-World Temporal Perception

Sep 2026 · 0 citations · 172 references
Computer Science

TL;DR

This paper articulates an alternative and complementary paradigm: perception and synthesis as distinct conditionings within a shared Generalist Open-World Temporal Perception Architecture (GOWTPA), positioning generalist temporal perception as a potential foundation layer for broader multimodal intelligence and physical AI.

Abstract

The next generation of artificial intelligence systems will likely be natively temporal and multimodal in both inputs and outputs: able to converse, perceive, predict, reason, and synthesize through a shared world representation. Realizing this requires a temporal perceptual substrate integrating sensory streams, language, and structured outputs within a multimodal world model. We seek a generalist open-world perceptual system that represents biological forms, natural physical structures, and artifacts, and their interactions, as a coherent, temporally persistent process. The model should infer geometry, articulation, semantics, interaction structure, and uncertainty from raw multimodal streams; maintain identity through occlusion and viewpoint change; generalize across species, forms, mechanisms, and materials; and abstain or expand its ontology when encountering the unknown. The objective is a structured world state supporting understanding, prediction, counterfactual reasoning, and controllable synthesis. Recent work suggests that some cross-modal and reasoning-like capabilities can emerge from large-scale generative video pretraining, reminiscent of language-model scaling. Yet these capabilities are often accessed through language probes or expressed through photorealistic video, leaving explicit semantic, geometric, or temporal structure largely unexposed. This paper articulates an alternative and complementary paradigm: perception and synthesis as distinct conditionings within a shared Generalist Open-World Temporal Perception Architecture (GOWTPA). Recognition, structured prediction, and simulation arise as different conditionings of the same generative substrate, while reasoning and embodiment-specific policies build upon the resulting world state. This positions generalist temporal perception as a potential foundation layer for broader multimodal intelligence and physical AI.

View source

Similar papers

Review Sep 2026

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models

For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world models has significantly enhanced the capabilities of robot learning systems, enab...

Shaunak A. Mehta, Ananya Hazarika, Haochen Zhang et al. · 0 citations
Aug 2026

Towards a Causally-inspired Evolving World Model for Vision-and-Language Navigation in Continuous Environments.

A causally-inspired evolving world modeling framework that formulates VLN-CE as a sequence of causal partially observable Markov decision processes that learns unified latent states that integrate vision, language, and action, and strengthens generalization across diverse navigation contexts.

Xuan Yao, Junyu Gao, Chang-Sheng Xu · 0 citations
Preprint Oct 2026

UniWAM: Unified World-Action Model

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...

Wen-Xuan Song, Jia-Yi Chen, Jing-Bo Wang et al. · 1 citation
Open access 2026

Teaching LLMs to Infer Real-World Consequences through Embodied Semantic Grounding

Embodied Semantic Grounding (ESG) is proposed, a framework that equips LLMs with consequence-aware text representations that extends affordance-grounding with consequence-level representations of post-event environmental functionality.

Manaswi Kulahara, Khadija Parwez, Faisal Alhwikem et al. · 0 citations
Review Open access Aug 2026

Intent-Driven Embodied Artificial Intelligence

This Review synthesizes recent progress in Embodied AI and articulate Intent-Driven Embodied Artificial Intelligence (IDEAI) as a system-level organizing framework in which intent functions as an explicit, revisable, and verifiable mediating construct between human goals, environmental constraints, and agent behavior.

Nanning Zheng · 0 citations
Open access Aug 2026

Predictive processing as a scalable computational principle for embodied multitask intelligence

A scalable hierarchical multimodal recurrent neural network grounded in predictive processing under the free-energy principle, capable of directly integrating more than 30,000-dimensional visuo-proprioceptive inputs without dimensionality reduction or handcrafted preprocessing is introduced.

Hayato Idei, Tamon Miyake, Tetsuya Ogata et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.