A Review of Multi-view 3D Reconstruction: From Classical Geometry to Feed-forward Models
Abstract
Multi-view 3D reconstruction undergoes several paradigm shifts over the past decades. This review categorizes the field into five representative paradigms: geometry-based Structure from Motion, learning-based Multi-view Stereo, Neural Radiance Fields, 3D Gaussian Splatting, and recent feed-forward geometric models. For each paradigm, we analyze its representation, key methods, advantages, and limitations, highlighting a clear transition from explicit geometric optimization to implicit neural representations, and further to efficient explicit modeling with pretrained feed-forward inference. Despite significant progress, challenges remain in geometric accuracy, rendering fidelity, and pose robustness. Future work is likely to focus on hybrid frameworks that combine geometric constraints, learned priors, explicit representations, and feed-forward inference to improve accuracy, efficiency, and generalization.