Skip to content

In-Training Masked Reconstruction as Structured Representation Augmentation for Collaboration-Aware V2X Perception

Sep 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 10999-11006 · 0 citations · 35 references

Abstract

LiDAR-based collaborative perception can mitigate occlusions by exchanging complementary viewpoints via Vehicle-to-Everything (V2X) communication. However, existing methods often depend heavily 3D annotations and adopt a pretraining pipeline that reconstruction serves as initialization. This letter presents a unified masked training framework that embeds online MAE reconstruction into multi-agent collaboration as a task-aware, structured representation augmentation, while jointly optimizing the downstream objective. Specifically, we first customize an ego-guided cross-agent masking strategy to explicitly exploit collaborative redundancy across agents. Then, a shared single-view backbone jointly encodes visible points with masked tokens, and an intermediate fusion module aggregates multi-agent features into a unified BEV representation. To stabilize joint optimization and improve convergence, we design a dual-head decoder. The reconstruction head predicts masked point distributions for both individual agents and collaborative scene, while the detection head performs task directly on the fused BEV. Besides, a distillation-inspired global-local consistency constraint aligns agent-level reconstructions with collaboration, preserving task-relevant semantics during training. Furthermore, our framework incurs no inference-time overhead beyond standard intermediate collaboration, and experiments on the V2X-Real benchmark demonstrate that it yields stronger collaborative BEV representations and achieves a +7.3% mAP$_{0.3}$ improvement on the vehicle-centric split over competitive baselines.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.