Speech-to-Plot: A Robust Voice-Driven Framework for Visual Grounding in Complex Acoustic Environments
In modern operational environments, rapid and hands-free target localization is crucial for situational awareness. However, traditional plotting systems rely on cumbersome manual interactions, and conventional multimodal algorithms degrade significantly under extreme background noise and constrained communication links. To address these challenges, we propose a novel edge-cloud collaborative Speech-to-Plot (STP) framework. The proposed system integrates a domain-adapted Automatic Speech Recognition (ASR) module—fine-tuned via a noise-injected curriculum—with a zero-shot visual grounding model to translate natural voice commands into precise spatial bounding boxes. Evaluations on a custom domain-specific dataset demonstrate that our framework exhibits graceful degradation rather than severe degradation under extreme acoustic interference, maintaining robust target semantic extraction even at 0 dB Signal-to-Noise Ratio. This reliable acoustic front-end helps prevent cascading errors in downstream cross-modal attention mechanisms, enabling accurate visual target localization. Furthermore, stress testing under simulated narrowband communication networks validates the practical engineering viability of our decoupled architecture. By offloading heavy multimodal inference to the cloud, the system mitigates computational congestion, supporting operational resilience despite the inevitable physical bandwidth limitations of field deployments.