Asynchrony-Robust Cooperative Perception and Prediction via Continuous-Time Global State Evolution
Abstract
Vehicle-to-everything (V2X) collaboration can alleviate the limited perception range and occlusion issues of single-agent autonomous driving. However, most existing cooperative studies still focus on single-frame perception, while the few works on joint cooperative perception and prediction largely rely on fixed-step, frame-aligned fusion, making them difficult to apply in realistic systems with cross-agent temporal misalignment. To address this issue, we propose CoSPACE, a novel framework that formulates asynchronous cooperative perception and prediction as continuous evolution and event-triggered correction of a shared scene state. Specifically, CoSPACE maintains an ego-centric global BEV state and propagates it to arbitrary observation timestamps using an ODE-style dynamics model, thereby preserving temporal coherence under irregular multi-agent inputs. When asynchronous observations arrive, they are treated as local evidence and assimilated by an event-triggered update module, which first selects informative regions and then performs time-aware gated correction to adaptively incorporate temporally reliable evidence. With this design, CoSPACE moves beyond the conventional frame-aligned fusion paradigm and mitigates the spatial misalignment, semantic confusion, and prediction degradation caused by temporal asynchrony. Experiments on V2XPnP-Seq under diverse delay settings show that CoSPACE consistently outperforms representative baselines and achieves strong robustness in asynchronous cooperative perception and prediction.