Beyond Euclidean Tokens: Hyperbolic Structure-Aware Mapping for Dual-Task Scene Parsing With Only Minimal Trainable Parameters.
Achieving unified scene parsing that simultaneously outputs cross-domain semantic segmentation and depth estimation without scene-specific retraining is crucial for robust perception in complex real-world environments, yet remains a challenging goal. While recent monocular depth estimation models such as DepthAnything...