Skip to content
Preprint

Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives

Aug 2026 · 0 citations · 60 references
Computer Science

TL;DR

This work presents a conceptual framework for Collaborative Joint Perception and Prediction (Co-P&P) that improves motion prediction of surrounding road users, thereby enhancing situational awareness in complex and dynamic traffic environments.

Abstract

Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions. We present a conceptual framework for Collaborative Joint Perception and Prediction (Co-P&P) that improves motion prediction of surrounding road users, thereby enhancing situational awareness in complex and dynamic traffic environments. Building upon our preliminary study, this extended version compares the performance of different fusion strategies and establishes baseline performance for a modular design of perception and prediction. Experimental results show that prediction-level fusion leads to a decline in overall system performance compared to detection-level or tracking-level fusion. We further implement a minimal end-to-end Co-P&P prototype that couples collaborative point-cloud sharing via the RENO neural codec with joint detection-forecasting via FutureDet, showing that collaboration improves forecasting accuracy while neural compression preserves this benefit at roughly 34x lower communication bandwidth.

View source

Similar papers

Oct 2026

Asynchrony-Robust Cooperative Perception and Prediction via Continuous-Time Global State Evolution

Vehicle-to-everything (V2X) collaboration can alleviate the limited perception range and occlusion issues of single-agent autonomous driving. However, most existing cooperative studies still focus on single-frame perception, while the few works on joint cooperative perception and prediction largely rely on fixed-step,...

Han-Xiao Ren, Ke-Qiang Li, Xiang Zhao et al. · 0 citations
Open access Aug 2026

Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA

Highlights What are the main findings? A latent-space LM-JEPA framework enables resource-efficient multi-modal object detection and scene perception for connected and autonomous vehicles, achieving higher perception accuracy with lower inference latency compared to conventional LLM and VLM-based methods. Context-aware...

Abhishek Gupta, Ajmery Sultana · 0 citations
Open access Aug 2026

Which2comm: An Efficient Collaborative Perception Framework with Connected and Automated Vehicles

Collaborative perception allows real-time inter-agent information exchange and thus offers invaluable opportunities to enhance the perception capabilities of individual agents. However, limited communication bandwidth in practical scenarios restricts the inter-agent data transmission volume. This implies a trade-of...

Duanrui Yu, Anqi Qu, Jing You et al. · 0 citations
Preprint Aug 2026

CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors

This paper proposes CoAnchor, an anchor-centric spatio-temporal alignment framework for asynchronous collaborative perception that builds sparse object-level spatio-temporal anchors as a shared interface for pose correction and tightly connects spatial refinement, temporal propagation, and current-time verification wit...

Chirong Li, Ruifeng Lin, Aobo Ji et al. · 0 citations
Preprint Aug 2026

Roadside-Cooperative Autonomous Driving: From Data Platform to Vision-Language End-to-End Reasoning

Vehicle-to-Everything (V2X) cooperation enables beyond-line-of-sight perception, mitigating occlusions in single-vehicle sensing. However, existing V2X benchmarks provide limited support for closed-loop evaluation and language-grounded supervision, hindering the development of vision-language models (VLMs) for end-to-e...

Yitao Xu, Tong Wu, Yi-Yang Wu et al. · 2 citations
Open access Sep 2026

CoSafe: A Cooperative V2V Perception Framework with LLM Reasoning for Hazard Detection on Real Dashcam Data

Cooperative perception through vehicle-to-vehicle (V2V) communication can resolve occlusions that single-vehicle systems cannot overcome, yet existing frameworks rely on simulated environments and expensive multi-sensor platforms. This paper presents CoSafe, a cooperative perception and reasoning framework built on rea...

Iosif-Alin Beti, P. Herghelegiu, C. Căruntu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.