Skip to content
Preprint

GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

Aug 2026 · 0 citations · 39 references
Computer Science

Abstract

Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions compose into goal-directed activities. Existing 4D scene graphs preserve object and place histories but omit activity structure, whereas activity representations are either not grounded in persistent 3D scenes or rely on externally provided event boundaries and object associations. We present GESTO (Grounded Event and Spatio-Temporal memOry), a spatio-temporal memory that couples a persistent 4D scene graph with a two-level hierarchy of atomic human--object interactions and goal-driven events. From an RGB-D observation stream, GESTO automatically extracts timestamped interactions, grounds them to persistent scene entities, groups them into events, and uses event context to refine uncertain object associations. A relation-aware tool-calling agent queries the resulting memory for activity-centric spatio-temporal reasoning. We evaluate GESTO on the reproducible text, binary, and time categories of an existing benchmark, together with 40 new Space2Event and Event2Space queries. GESTO achieves scores of 0.71, 0.75, and 0.70 on the standard categories, approaching a method supplied with ground-truth event and object grounding, while substantially outperforming the same reasoning framework when these inputs are removed. It further achieves 0.73 and 0.75 on Space2Event and Event2Space queries. Ablations show that hierarchical event structure and context-aware grounding refinement provide complementary benefits, supporting activity-grounded hierarchical memory for retrospective reasoning in dynamic human environments.

View source

Similar papers

Jun 2026

Analytic Concept-Centric Memory for Agentic Embodied Manipulation

This work proposes an analytic concept-centric memory framework for agentic embodied manipulation that improves task completion, retrieval accuracy, object re-identification, and cross-object skill generalization over unstructured and embedding-based memory baselines.

Mingyang Sun, Xiujian Liang, Jiude Wei et al. · 0 citations
Open access Jul 2026

Symbolic-driven agentic reasoning for environmental and behavioral event detection

SDAR employs a symbolic reasoning engine to guide agentic decision making, connecting low-level visual cues with structured symbolic representations of events, and enables interpretable reasoning chains that capture causal relationships, contextual dependencies, and event categories.

Guangyao Chen, Liqin Luo, Jun Peng et al. · 0 citations
#natural language process... Preprint Jul 2026

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

The overall results show that editable, consolidated memory can supply remembered context for robot planning, and full MEMORA--combining editing, typed stores, and consolidation--achieves the strongest aggregate results among the evaluated memory conditions.

Zihao Yu, Xiu Yuan, Chongjie Zhang · 0 citations
Preprint Aug 2026

SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation

This work presents SuperMap, a 4D spatio-temporal mapping framework for language-guided navigation that integrates high-frequency geometric SLAM with asynchronous open-vocabulary perception and releases the full system as open-source to provide the community with a deployable baseline for open-vocabulary spatio-temporal mapping.

Shibo Zhao, Guofei Chen, Honghao Zhu et al. · 1 citation
Preprint Jul 2026

Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs

SG-Ego, a large scale annotation set extending Ego4D with spatio-temporal scene graphs, where relations triplets are consolidated over time into explicit time-evolving descriptions of the scene state is introduced, and GLEN, a graph-based model that operates over scene graph sequences to both align them with textual actions and model their temporal evolution is proposed.

Francesca Pistilli, Simone Alberto Peirone, Giuseppe Averta · 0 citations
Preprint Aug 2026

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

LT-Mem is proposed, a volatility-aware memory evolution framework that unifies spatially aligned instance-level 3D perception with volatility-conditioned temporal reasoning and introduces LT-VQA, a dataset and evaluation suite comprising multi-session recordings, persistent identity annotations, and temporal QA pairs.

Yumin Lee, Hyoseok Ju, Giseop Kim · 0 citations