Skip to content

Author

Ajmery Sultana

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA

Highlights What are the main findings? A latent-space LM-JEPA framework enables resource-efficient multi-modal object detection and scene perception for connected and autonomous vehicles, achieving higher perception accuracy with lower inference latency compared to conventional LLM and VLM-based methods. Context-aware and adaptive sensor fusion, selective latent transmission, and lightweight edge-assisted reasoning improve cooperative scene understanding, yielding up to 25% better scene understanding, 20% higher intersection success rates, and a 15% reduction in transmitted model parameters. What are the implications of the main findings? Latent representation learning provides a practical alternative to token-based LLM and VLM inference, making real-time multi-modal perception feasible on resource-constrained edge devices in connected and autonomous vehicles. The proposed framework demonstrates that adaptive latent communication and collaborative reasoning can enhance the scalability, energy efficiency, and safety of future intelligent transportation and cooperative autonomous driving systems. Abstract This paper presents the latent model-joint embedding predictive architecture (LM-JEPA), a resource-efficient collaborative perception framework for connected and autonomous vehicles that integrates latent predictive representation learning with lightweight multi-modal reasoning. Autonomous driving in urban and highway environments requires accurate scene understanding under strict latency, energy, and communication constraints, limiting the practicality of large language model (LLM) and vision–language model (VLM)-based approaches in edge deployments. To address this, LM-JEPA encodes heterogeneous inputs including camera, LiDAR, radar, and map data into a unified latent space using a joint embedding predictive architecture, enabling efficient perception and reasoning without token-level inference. Unlike existing latent-space learning approaches that primarily learn predictive visual embeddings for single-modal perception, the proposed framework integrates multi-modal latent reasoning and adaptive sensor fusion to support collaborative perception under resource-constrained vehicular edge environments. The collaborative perception framework introduces a context-adaptive multi-modal fusion mechanism that dynamically weights sensor and model contributions, along with selective latent transmission and adaptive decoding for resource-aware operation. A lightweight VLM is integrated with an edge-assisted vehicular pipeline to support real-time on-vehicle inference with adaptive offloading based on latency and energy constraints, while a latent-space reasoning module enables cooperative decision-making. Experiments on BDD100K and nuScenes-QA show that LM-JEPA improves perception accuracy by 5% and reduces latency by approximately 7% over LLM and VLM baselines, while achieving up to 25% improvement in scene understanding, 20% higher intersection success rates, improved highway merging, and approximately 15% reduction in the transmitted model parameters.

Abhishek Gupta, Ajmery Sultana · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.