Conference
Aug 2026
ActionLMM: captioning long-video actions with memory-augmented VLMs
This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.
Ruirui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al.
· International Conference on... · 0 citations