DOpt: Decomposed Optimization for Collaborative VLM Inference Over IoT-UAV-GBS Hierarchies
Abstract
Vision-language models (VLMs) have achieved remarkable success, yet their substantial resource demands far exceed the capabilities of typical IoT devices. This paper investigates collaborative VLM inference across IoT devices, mobile UAV relays, and a ground base station to bring intelligence to the edge. The collaborative inference problem can be formulated as a constrained optimization that minimizes a weighted sum of delay, energy, and inference distortion over time. However, a direct solution is challenging due to rapidly changing wireless channels, coupled with per-slot resource constraints across multiple tasks, and the presence of discrete decision variables. Our core idea is to decouple the problem: we propose DOpt, a framework in which slow UAV mobility is learned via multi-agent reinforcement learning and fast per-slot compression allocation is solved via convex optimization. Extensive simulations demonstrate that DOpt improves the weighted objective by up to 23.1% over all baselines, with ablation studies confirming the necessity of both adaptive compression and learned mobility.