A Timeliness-Aware Large-Small VLM Collaboration (TALSC) framework that first model the Age of Information (AoI) evolution for VLM inference and characterize the coupling among AoI, token length, and task performance to formulate a general timeliness metric and proposes the TALSC online scheduling algorithm.
Abstract
The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and reasoning capabilities. Infrastructure-assisted AD alleviates this resource constraint by enabling collaboration with large VLMs (LVLMs) at edge servers. However, in dynamic vehicular environments, the utility of sensory data for downstream tasks decays rapidly, making timeliness of information a critical concern. To balance the accuracy gains of LVLMs with their latency-induced timeliness degradation, we develop a Timeliness-Aware Large-Small VLM Collaboration (TALSC) framework. Specifically, we first model the Age of Information (AoI) evolution for VLM inference and characterize the coupling among AoI, token length, and task performance to formulate a general timeliness metric. Building on this, we propose the TALSC online scheduling algorithm. Since scheduling decisions have a delayed impact on future timeliness metric and the output token number is unknown at scheduling time, we design a Lyapunov drift-plus-estimated-penalty algorithm and provides a guaranteed performance. In simulation, we first conduct a case study to derive a fitted timeliness metric based on nuScenes dataset, and further show that TALSC outperforms baselines under various communication and computing settings, achieving up to a 12.6\% normalized improvement in Micro-F1 score compared with the best-performing baseline.
ComVLA is proposed, a framework that uses the dense semantic information contained in the language guidance to adapt the VLA token budget to the channel capacity, and demonstrates that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.
Bo-Liang Liu, Wint Yi Poe, Jing-Yun Di et al.· 0 citations
Vision-Language-Action (VLA) models are reshaping autonomous driving (AD) by unifying perception, reasoning, and control through language, enabling semantic grounding, interpretable decisions, and better long-tail generalization. But language is expensive onboard: latency and memory budgets are tight, and autoregressiv...
Vehicle-to-everything (V2X) techniques expand the capability boundaries of connected and autonomous vehicles (CAVs). However, deploying V2X-enabled end-to-end autonomous driving (E2E-AD) systems still faces a trade-off between sharing high-resolution perception features and V2X communication bandwidth constraints. Furt...
Han Jiang, Zi-Yang Yan, Zechang Ye et al.· IEEE Transactions on Cogniti...· 0 citations
A risk-adaptive edge-cloud architecture in which onboard traffic assessment determines when cloud reasoning is requested is presented, in which onboard traffic assessment served as a practical trigger for selective VLM inference in these experiments.
Meng Ma, Shu-Yang Li, Nai-Gang Wang et al.· 0 citations
Combining two-step denoising with representative execution methods substantially reduces inference cost with a small reduction in task performance, motivating joint optimization of model-inference efficiency and robot-system timing.
Di Wu, Rong-Tian Shen, Ping Liu et al.· 0 citations
Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavio...
Rithvik Jonna, Man Namgung, Aakash Gurram et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.