Skip to content
Open access

Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation

Aug 2026 · International Journal of Computer Vision · Vol 134 · 0 citations · 44 references
Computer Science

TL;DR

A new top-down paradigm is proposed that maintains a unified, instance-centric human-aware 3D space, enabling simultaneous camera calibration, cross-view association, and human reconstruction via cross-modal contrastive learning and allows correspondence reasoning, semantic aggregation, and instance discrimination to be performed natively in the 3D space.

Abstract

Multi-view human reconstruction has been extensively studied under simplified settings, yet scaling these methods to robust and efficient multi-person reconstruction in unconstrained environments requires a more general and scalable modeling paradigm. Existing bottom-up methods often rely on accurate camera calibration and explicit cross-view matching, and therefore struggle in multi-person scenarios with severe occlusions and ambiguities. We attribute these limitations to the lack of an explicit and view-agnostic 3D representation for jointly reasoning about human structure and multi-view consistency. Based on this observation, we propose a new top-down paradigm that maintains a unified, instance-centric human-aware 3D space, enabling simultaneous camera calibration, cross-view association, and human reconstruction via cross-modal contrastive learning. Specifically, observations from multiple views are lifted and fused directly into this shared 3D space, where geometric structure, visual appearance, and human-centric semantic cues are jointly encoded at the instance level. To enhance the discriminability and consistency of the 3D representation, we introduce a spatial contrastive learning strategy that aligns 3D features corresponding to the same human instance across different views and modalities, while separating features from different instances. This design allows correspondence reasoning, semantic aggregation, and instance discrimination to be performed natively in the 3D space, which enforces cross-view consistency and improves robustness under severe occlusions. Based on the learned instance-aware 3D representation, we recover structured human body models in a feed-forward manner by regressing SMPL parameters from instance-level 3D human tokens. Extensive experiments demonstrate that adopting a unified human-aware 3D feature space as the core representation leads to robust, accurate, and efficient multi-view human reconstruction in challenging real-world scenarios. The code and data will be made publicly available.

Read PDF

Similar papers

Preprint Sep 2026

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied...

Ao-Xiang Fan, Corentin Dumery, Nicolas Talabot et al. · 0 citations
Conference Aug 2026

Multiview Consistency Learning with Synthetic Camera Views for Hand-in-the-Wild Gesture Recognition

Hand gesture recognition is an important problem in human–computer interaction, virtual reality, augmented reality, robotics, and multimedia systems. Despite strong performance in previous studies, recognition in uncontrolled environments remains challenging due to viewpoint variation, self-occlusion, articulation ambi...

M. Nguyen, Hai Vu, Huong-Giang Doan et al. · 0 citations
Preprint Sep 2026

An Evaluation Framework for Generating Multi-View Images of a Person in a Scene

The Head Scene Rotation Difference (HSRD) metric is proposed to quantitatively evaluate camera movements around a person, providing the pipeline necessary to evaluate 3D spatial parallax for a person in a scene, paving the way to reliably construct high-quality multi-view synthetic datasets.

M. Majid, Young Kyung Kim, Guillermo Sapiro · 0 citations
#machine learning Preprint Sep 2026

SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation

Consistent multi-view object segmentation is critical for 3D perception and robotics, yet remains challenging under severe viewpoint and occlusion changes. Existing methods typically perform 3D instance segmentation on point clouds or rely on offline 2D mask-matching pipelines. However, 3D instance segmentation is limi...

Jiang-Shan Gong, Yuqun Wu, Qiqian Fu et al. · 1 citation
Preprint Aug 2026

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion stude...

Jiakun Li, Li Fang, Hao Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.