Skip to content
#edge computing Preprint

ComVLA: Communication-Aware Split Inference for VLA Models in 6G-Connected Robotics

Sep 2026 · 0 citations · 18 references
Computer Science

TL;DR

ComVLA is proposed, a framework that uses the dense semantic information contained in the language guidance to adapt the VLA token budget to the channel capacity, and demonstrates that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.

Abstract

Connected robotics is an emerging 6G application where mobile robots follow natural-language instructions to manipulate physical objects. The Vision-Language-Action (VLA) models that enable this are too large to run on the robot; a common trend is to offload inference to the cloud. The wireless link, however, limits how much sensing data the edge can transmit per control step. Two recent lines address this constraint: semantic communication codecs compress sensor data but require channel-specific retraining, and VLA token pruners select tokens from image but ignore the channel. Our insight is that the dense semantic information contained in the language already indicates which visual tokens matter. We propose ComVLA, a framework that uses this language guidance to adapt the VLA token budget to the channel capacity. Transmitting 32 tokens instead of 512 on the LIBERO benchmark, ComVLA cuts inference compute by 74% and inference latency by 22% versus the original OpenVLA-OFT baseline, at a cost of 1.5 pp in average task success (95.4% vs. 96.9%), and it stays within the capacity budget under Rayleigh and Rician fading. These results demonstrate that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.

View source

Similar papers

Preprint Sep 2026

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Robo-Harness K1 is introduced, a robot-use agent (RUA) framework that exposes perception as tools that makes 3D geometry accessible without changing the VLM architecture or training a depth encoder, and suggests that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies t...

Ze-Xi Li, Ye-Hang Zhang, Wen-Qian Li et al. · 2 citations
#machine learning Preprint Sep 2026

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

This work proposes a framework that exposes backbone depth V, action expert depth A, and denoising steps $D$ as three jointly configurable compute axes in a VLA, and introduces a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deep...

Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro et al. · 0 citations
Sep 2026

VLM-Driven Robotic Control in AI-RAN: System Design and Task-Aware Optimization

AI-native radio access networks (AI-RAN) are evolving from connectivity infrastructures into edge execution platforms. Robotic control driven by vision-language models (VLMs) is a representative application scenario, where limited onboard computing capability often requires visual data to be uploaded to the edge server...

Ling-Xiao Sun, Zhao-Yang Zhang, Zi-Rui Chen et al. · 0 citations
Preprint Sep 2026

EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model

Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavio...

Rithvik Jonna, Man Namgung, Aakash Gurram et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Show-Harness: Just a VLM Agent Can Play Robots

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to"play"robots through a compact semantic interface linking intent to action. Show-Harness exposes...

Yan-Zhe Chen, Ze-Chen Bai, Zhi-Jun Cao et al. · 23 citations · ⚡2
Preprint Sep 2026

Goal-Oriented Communications for Physical AI: Design and Testbed

Physical AI relies on frequently-updated, latency-sensitive video stream to perceive, reason, and interact with the physical world, resulting in strict latency requirements with much higher data volumes that existing 5G networks cannot support. Goal-oriented communication (GoC) offers as a promising approach to solve t...

Shu-Tong Chen, Wen-Kai Zhang, Adnan Aijaz et al. · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.