VLM-Driven Robotic Control in AI-RAN: System Design and Task-Aware Optimization
Abstract
AI-native radio access networks (AI-RAN) are evolving from connectivity infrastructures into edge execution platforms. Robotic control driven by vision-language models (VLMs) is a representative application scenario, where limited onboard computing capability often requires visual data to be uploaded to the edge server for VLM inference. Higher visual quality may provide richer semantic information, but it can also increase uplink load and inference latency. This paper proposes a task-aware video adaptation mechanism for adapting region-of-interest (ROI) video streams to address the latencysemantic tradeoff under dynamic wireless conditions. The ROI video quality level is treated as an edge-side control variable that affects both communication and inference latency, while determining the task-relevant semantic information available to the VLM. Since each quality decision affects subsequent latency and semantic feedback, the online adaptation process naturally becomes a sequential decision-making problem. We formulate it as a Markov decision process (MDP) and design a lightweight Q-learning controller to meet latency service-level objectives. Experiments on a real 5G edge robotic prototype validate the effectiveness of the proposed scheme. Compared with a rule-based baseline, the proposed controller reduces the 95th-percentile end-to-end latency by 21.3% and improves the average task-relevant semantic availability by $\mathbf{1 2. 2\%}$.