This work proposes an e-ective pre-training strategy, namely Temporal Masked Auto-Encoders (T-MAE), which takes as input temporally adjacent frames and learns temporal dependency, and demonstrates that T-MAE achieves the best performance on both Waymo and ONCE datasets among competitive self-supervised approaches.
Vernata is introduced, consisting of three extensions: sparse view augmentation to improve robustness against varying point densities, a memory bank mechanism to stabilize resource-constrained training, and cross-modal distillation utilizing dense, high-resolution 2D image features to enable fine-grained semantic guida...
Oliver Lemke, Alexander Liniger, Abel Gawel et al.· 0 citations
A SAM-guided framework for point cloud oversegmentation that significantly improves boundary recall and maintains high oracle accuracy while maintaining high oracle accuracy, and generalizes well to unseen datasets without retraining, showing strong cross-dataset inference capability.
Dening Lu, Michael A. Chapman, Jonathan Li· The International Archives o...· 0 citations
This work proposes a cascade optimization framework that systematically enhances feature representation and refines multimodal fusion, and introduces the Multi-Scale Contextual Fusion Module (MSCF) to reduce alignment bias.
Three-dimensional object detection from LiDAR point clouds is essential for autonomous driving, yet existing methods typically rely on costly and extensively annotated datasets. Self-supervised masked autoencoders (MAE) provide a promising alternative, but effectively capturing local geometric structures while maintain...
Jia-Yi Zhou, Yaqian Ning, Jie Cao et al.· Remote Sensing· 0 citations
A zero-shot, open-vocabulary semantic segmentation framework for ALS point clouds based on 2D-3D transfer, utilizing three types of VFMs and introduces an adaptive global view projection module that derives optimal virtual camera poses and field-of-view (FOV) from scene extents, effectively enabling the application of...
Yang-Hong Lin, Tian-Yu Li, Shu-Dong Zhou et al.· The International Archives o...· 1 citation
Four key contributions are highlighted: SAC for broader spatial contextual relationships, LFR for improved feature preservation, their integration into an efficient segmentation framework, and the introduction of two sparse point cloud-derived road-marking segmentation datasets, which are publicly available at this lin...
M. L. R. Lagahit, Xin Liu, Haoyi Xiu et al.· IEEE Journal of Selected Top...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.