Aug 2026· International Journal of Computer Vision· Vol 134· 0 citations· 44 references
Computer Science
TL;DR
A new top-down paradigm is proposed that maintains a unified, instance-centric human-aware 3D space, enabling simultaneous camera calibration, cross-view association, and human reconstruction via cross-modal contrastive learning and allows correspondence reasoning, semantic aggregation, and instance discrimination to be performed natively in the 3D space.
Abstract
Multi-view human reconstruction has been extensively studied under simplified settings, yet scaling these methods to robust and efficient multi-person reconstruction in unconstrained environments requires a more general and scalable modeling paradigm. Existing bottom-up methods often rely on accurate camera calibration and explicit cross-view matching, and therefore struggle in multi-person scenarios with severe occlusions and ambiguities. We attribute these limitations to the lack of an explicit and view-agnostic 3D representation for jointly reasoning about human structure and multi-view consistency. Based on this observation, we propose a new top-down paradigm that maintains a unified, instance-centric human-aware 3D space, enabling simultaneous camera calibration, cross-view association, and human reconstruction via cross-modal contrastive learning. Specifically, observations from multiple views are lifted and fused directly into this shared 3D space, where geometric structure, visual appearance, and human-centric semantic cues are jointly encoded at the instance level. To enhance the discriminability and consistency of the 3D representation, we introduce a spatial contrastive learning strategy that aligns 3D features corresponding to the same human instance across different views and modalities, while separating features from different instances. This design allows correspondence reasoning, semantic aggregation, and instance discrimination to be performed natively in the 3D space, which enforces cross-view consistency and improves robustness under severe occlusions. Based on the learned instance-aware 3D representation, we recover structured human body models in a feed-forward manner by regressing SMPL parameters from instance-level 3D human tokens. Extensive experiments demonstrate that adopting a unified human-aware 3D feature space as the core representation leads to robust, accurate, and efficient multi-view human reconstruction in challenging real-world scenarios. The code and data will be made publicly available.
Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied...
Ao-Xiang Fan, Corentin Dumery, Nicolas Talabot et al.· 0 citations
Hand gesture recognition is an important problem in human–computer interaction, virtual reality, augmented reality, robotics, and multimedia systems. Despite strong performance in previous studies, recognition in uncontrolled environments remains challenging due to viewpoint variation, self-occlusion, articulation ambi...
M. Nguyen, Hai Vu, Huong-Giang Doan et al.· International Conference on...· 0 citations
The Head Scene Rotation Difference (HSRD) metric is proposed to quantitatively evaluate camera movements around a person, providing the pipeline necessary to evaluate 3D spatial parallax for a person in a scene, paving the way to reliably construct high-quality multi-view synthetic datasets.
M. Majid, Young Kyung Kim, Guillermo Sapiro· 0 citations
Consistent multi-view object segmentation is critical for 3D perception and robotics, yet remains challenging under severe viewpoint and occlusion changes. Existing methods typically perform 3D instance segmentation on point clouds or rely on offline 2D mask-matching pipelines. However, 3D instance segmentation is limi...
Jiang-Shan Gong, Yuqun Wu, Qiqian Fu et al.· 1 citation
DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion stude...
Jiakun Li, Li Fang, Hao Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.