Skip to content

Pretraining Body Part Representations for Text-Motion Retrieval

Jul 2026 · ACM Transactions on Multimedia Computing, Communications, and Applications (TOMCCAP) · 0 citations · 65 references

TL;DR

This work proposes POP-TMR which pretrains body part representations for fine-grained text-motion retrieval and introduces HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment.

Abstract

Text-motion retrieval has gained increasing research attention, yet several critical challenges remain such as data scarcity, limited fine-grained matching capabilities, and inadequate evaluation protocols. To address these issues, we propose POP-TMR which pretrains body part representations for fine-grained text-motion retrieval. Our approach leverages large-scale human motion datasets to pretrain a spatio-temporal transformer-based motion encoder, enabling more generalizable motion features. In addition to matching global motion and text representations, we propose a local branch to capture detailed body part features for enhancing spatial-aware cross-modal alignment. To improve evaluation, we introduce HumanML3D+, an enhanced benchmark that provides accurate positive annotations for text queries and includes text descriptions at varying levels of detail, enabling more systematic performance assessment. Extensive experiments on KIT-ML, HumanML3D and HumanML3D+ benchmarks demonstrate that POP-TMR outperforms state-of-the-art methods. Furthermore, we showcase its effectiveness in additional downstream applications, including text-to-motion generation evaluation, human interaction recognition and zero-shot moment retrieval. Data, code and pretrained model are publicly available at https://lin-kayla.github.io/POP_TMR/.

View source

Similar papers

Preprint Aug 2026

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

This work proposes a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model that improves mixed-granularity retrieval without compromising standard-caption performance, and believes that its MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.

Fulong Liu, Liang Xu, Chengqun Yang et al. · 0 citations
Preprint Aug 2026

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form descriptions typically provide supervision only at the clip level, without explicit temporal correspondence between motion frames and language. This limits fine-grained motion--text grounding and temporally precise generation. We propose FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations. Our method first segments long-form descriptions into action-bearing phrases, and then formulates motion--language alignment as an optimal transport problem, which naturally models many-to-many relations between motion frames and text under global constraints. With entropic regularization and Sinkhorn iterations, FineMoLA efficiently infers pseudo frame-level alignments without human labeling. Experiments on SnapMoGen demonstrate that the learned alignments outperform baselines in motion--text grounding.

Tongyan Wang, Zhengyuan Li, Muhan Lin et al. · 0 citations
Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Ruirui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations
Book Open access Jul 2026

Generation-Augmented Video Corpus Moment Retrieval

This work proposes Video-GAR, a novel framework that reframes the conventional retrieval task from superficial matching to generative understanding, positing that the capability for query reconstruction evidences deep semantic comprehension.

Mingjin Kuai, Qianyin Xiao, Juncheng Li et al. · 0 citations
Preprint Jul 2026

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

This work proposes Self-SiMS, a self-similarity-based Moment Proposal and Scoring that exploits intrinsic relationships within videos, enabling robust span generation and scoring and introduces a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video.

Jihyun Lee, Cheol-Ho Cho, Woojin Jun et al. · 0 citations