Beyond Euclidean Tokens: Hyperbolic Structure-Aware Mapping for Dual-Task Scene Parsing With Only Minimal Trainable Parameters.
Achieving unified scene parsing that simultaneously outputs cross-domain semantic segmentation and depth estimation without scene-specific retraining is crucial for robust perception in complex real-world environments, yet remains a challenging goal. While recent monocular depth estimation models such as DepthAnything V2 exhibit strong domain generalization, semantic segmentation still suffers from severe structural degradation under domain and viewpoint shifts. We observed a persistent hierarchical calibration gap, where Euclidean representations exhibit larger calibration gaps between child and parent categories under domain shifts, suggesting limitations of existing Euclidean-based methods in preserving semantic hierarchies. To address this issue, we propose HyperMapper, a hyperbolic structure-aware mapping framework that bridges semantic understanding and geometric priors through hyperbolic token-to-feature interactions. By exploiting the negative curvature of hyperbolic space, HyperMapper helps capture hierarchical relationships and maintains geometric consistency across domains. Furthermore, by combining the expressive priors of vision foundation models (VFMs) with parameter-efficient fine-tuning (PEFT), HyperMapper achieves cross-domain adaptation with minimal trainable parameters in backbone while retaining the strong depth estimation capability of DepthAnythingV2 without retraining. Extensive experiments on multiple cross-domain and cross-viewpoint benchmarks demonstrate that HyperMapper achieves a higher mIoU for both parent and child categories while consistently improving segmentation accuracy over strong baselines. Our approach establishes a promising direction for task-preserving dual-task adaptation, bridging semantic and geometric learning and paving the way toward unified, cross-domain scene parsing.