In-Training Masked Reconstruction as Structured Representation Augmentation for Collaboration-Aware V2X Perception
Abstract
LiDAR-based collaborative perception can mitigate occlusions by exchanging complementary viewpoints via Vehicle-to-Everything (V2X) communication. However, existing methods often depend heavily 3D annotations and adopt a pretraining pipeline that reconstruction serves as initialization. This letter presents a unified masked training framework that embeds online MAE reconstruction into multi-agent collaboration as a task-aware, structured representation augmentation, while jointly optimizing the downstream objective. Specifically, we first customize an ego-guided cross-agent masking strategy to explicitly exploit collaborative redundancy across agents. Then, a shared single-view backbone jointly encodes visible points with masked tokens, and an intermediate fusion module aggregates multi-agent features into a unified BEV representation. To stabilize joint optimization and improve convergence, we design a dual-head decoder. The reconstruction head predicts masked point distributions for both individual agents and collaborative scene, while the detection head performs task directly on the fused BEV. Besides, a distillation-inspired global-local consistency constraint aligns agent-level reconstructions with collaboration, preserving task-relevant semantics during training. Furthermore, our framework incurs no inference-time overhead beyond standard intermediate collaboration, and experiments on the V2X-Real benchmark demonstrate that it yields stronger collaborative BEV representations and achieves a +7.3% mAP$_{0.3}$ improvement on the vehicle-centric split over competitive baselines.