Skip to content

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

Sep 2026 · 0 citations
Computer Science

TL;DR

Kairos is introduced, a video dataset for video-language modeling with time-resolved annotations that supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation.

Abstract

Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. Kairos provides a general-purpose foundation for modeling visual experiences over time.

View source

Similar papers

Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations
#small language model Preprint Aug 2026

Training-Free Temporal Abstraction for General Video Understanding

STITCH is presented, a training-free method that divides a video into semantically meaningful temporal chunks that are computed once per video and reused across tasks, suggesting that reusable temporal abstraction is a promising direction for general video understanding.

Etienne Casanova, S. Brodjian, Pietro Perona · 0 citations
Preprint Aug 2026

Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

Video-language models (VLMs) remain brittle on tasks that require tracking events over time and grounding answers in specific spatial regions. We propose that part of this limitation can be addressed through better organization of visual evidence at inference time. We introduce structured video prompting, a training-fr...

Sadegh Mohammadian · 0 citations
#artificial intelligence Preprint Sep 2026

Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception

Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as timestamp tokens. However, this generative f...

Yan Shi, Yan Song, Jun-Hui Li et al. · 0 citations
Preprint Sep 2026

The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance.

Yu-Meng Shi, Quanyu Long, Yin Wu et al. · 0 citations
Preprint Aug 2026

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

VideoRouter (VR) is proposed that rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames, and introduces a verification-guided router to determine which view is better supported by the selected evidence and select the final answer.

Ziling Huang, Yuki M. Asano, Shin'ichi Satoh · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.