for Point Cloud Representation Learning
This work proposes an e-ective pre-training strategy, namely Temporal Masked Auto-Encoders (T-MAE), which takes as input temporally adjacent frames and learns temporal dependency, and demonstrates that T-MAE achieves the best performance on both Waymo and ONCE datasets among competitive self-supervised approaches.