Skip to content
Preprint

An Evaluation Framework for Generating Multi-View Images of a Person in a Scene

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

The Head Scene Rotation Difference (HSRD) metric is proposed to quantitatively evaluate camera movements around a person, providing the pipeline necessary to evaluate 3D spatial parallax for a person in a scene, paving the way to reliably construct high-quality multi-view synthetic datasets.

Abstract

Recent generative image-editing Diffusion Transformers (DiTs) demonstrate impressive semantic editing capabilities but still struggle with spatially consistent camera angle changes. A primary bottleneck in training foundation models to execute free-form, promptable camera angle changes is the lack of specialized training data. While multi-view datasets exist for generic 3D environments and objects, there remains an absence of paired, multi-view datasets featuring human subjects at fixed locations in natural scenes, including frontal and side-profile views. Capturing such multi-camera data in unconstrained environments is logistically challenging and unscalable. In this paper, we first experiment with multiple state-of-the-art image editing models to create this data synthetically, but find that the outputs are frequently prone to hallucinations involving how much the subject's head turns relative to the background, often producing inconsistent environments. To address this issue, we propose the Head Scene Rotation Difference (HSRD) metric to quantitatively evaluate camera movements around a person. The proposed metric operates by decoupling camera movement from localized head pose manipulation. As demonstrated by the extensive experimentation, HSRD provides the pipeline necessary to evaluate 3D spatial parallax for a person in a scene, paving the way to reliably construct high-quality multi-view synthetic datasets.

View source

Similar papers

Conference Aug 2026

Multiview Consistency Learning with Synthetic Camera Views for Hand-in-the-Wild Gesture Recognition

Hand gesture recognition is an important problem in human–computer interaction, virtual reality, augmented reality, robotics, and multimedia systems. Despite strong performance in previous studies, recognition in uncontrolled environments remains challenging due to viewpoint variation, self-occlusion, articulation ambi...

M. Nguyen, Hai Vu, Huong-Giang Doan et al. · 0 citations
Preprint Sep 2026

Revisiting Multi-View Stereo: A Sequence-to-Sequence Formulation

Computing accurate geometry from multi-view images is a fundamental problem in computer vision. Recent feed-forward (FF) models jointly estimate 3D geometry and camera parameters, but they typically suffer from geometry distortion caused by reconstruction ambiguity, even when ground-truth camera parameters are supplied...

Ao-Xiang Fan, Corentin Dumery, Nicolas Talabot et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Geometric Inconsistency Localization in Multi-View Image Sets

DeformView is introduced, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies and DEFECt3R is proposed, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level.

Xander Staelens, Albéric Loos, Bert Ramlot et al. · 0 citations
Open access Aug 2026

Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation

A new top-down paradigm is proposed that maintains a unified, instance-centric human-aware 3D space, enabling simultaneous camera calibration, cross-view association, and human reconstruction via cross-modal contrastive learning and allows correspondence reasoning, semantic aggregation, and instance discrimination to b...

Yuanwang Yang, Buzhen Huang, Zong-Xuan Ren et al. · 0 citations
Preprint Aug 2026

GenRec: Knowing Where to Reconstruct and Where to Generate

GenRec is introduced, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow, and attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved o...

Ata Çelen, Jaewoo Jung, Federico Tombari et al. · 0 citations
Preprint Aug 2026

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion stude...

Jiakun Li, Li Fang, Hao Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.