The Depth Segment Anything Model (DepthSAM), a MDE-adapted method specifically designed to mitigate this misalignment, achieves both robust semantic understanding in camouflaged environments and accurate segmentation of camouflaged objects.
A mask-aware tri-modal framework that improves the quality of superpoint representations by retrieving a scene-level structural context from a pretrained PointSAM encoder to enhance object-centric evidence and predicting a soft mask weight to suppress unreliable superpoints.
Feng Zhou, Hui Wang, Kaida Ning et al.· The Visual Computer· 0 citations
360{\deg} salient object detection (SOD) aims to accurately segment salient regions across a full field of view. However, equirectangular projection (ERP) introduces severe spatial distortion when mapping the spherical domain onto a planar representation. Existing methods mainly focus on compensating projection distort...
Jun-Song Zhang, Zhi-Jie Shen, Shuai Zheng et al.· 0 citations
This work enhances the existing iterative object-basesd visual localization approach with an additional semantic feature derived from a pretrained semantic segmentation model and conducts a systematic baseline study of contemporary feature matching techniques on such cross-domain query-reference image pairs.
Yasmin Loeper, Markus Gerke, P. Fanta-Jende· The International Archives o...· 0 citations
Marigold V2 achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition, and achieves state-of-the-art results when applied to other dense regression tasks, such as surface normals estimation and intrinsic image decomposition.
Igor Pavlovic, Thiemo Wandel, Anton Obukhov et al.· 0 citations
This work proposes three learned matching heads: a LightGlue-style attention head with DoubleSoftmax scoring on frozen MASt3R descriptors; a DPT-style multi-scale fusion module that exposes layered spatial detail from the VGGT foundation model before pooling; and a multi-view extension that performs joint self-attentio...
Denis Fatykhoph, Timur Akhtyamov, Konstantin Pakulev et al.· arXiv.org· 0 citations
High-precision perception is fundamental to safe autonomous driving, and BEV-based 3D object detection via lidar-camera fusion plays a crucial role in improving detection accuracy and robustness. To address insufficient feature representation, spatial misalignment, and the limitations of static fusion strategies, this...
Jie Hu, Xinghao Cheng, Shuaidi He et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.