Skip to content

Author

Qing-Xi Yang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Speech-to-Plot: A Robust Voice-Driven Framework for Visual Grounding in Complex Acoustic Environments

In modern operational environments, rapid and hands-free target localization is crucial for situational awareness. However, traditional plotting systems rely on cumbersome manual interactions, and conventional multimodal algorithms degrade significantly under extreme background noise and constrained communication links. To address these challenges, we propose a novel edge-cloud collaborative Speech-to-Plot (STP) framework. The proposed system integrates a domain-adapted Automatic Speech Recognition (ASR) module—fine-tuned via a noise-injected curriculum—with a zero-shot visual grounding model to translate natural voice commands into precise spatial bounding boxes. Evaluations on a custom domain-specific dataset demonstrate that our framework exhibits graceful degradation rather than severe degradation under extreme acoustic interference, maintaining robust target semantic extraction even at 0 dB Signal-to-Noise Ratio. This reliable acoustic front-end helps prevent cascading errors in downstream cross-modal attention mechanisms, enabling accurate visual target localization. Furthermore, stress testing under simulated narrowband communication networks validates the practical engineering viability of our decoupled architecture. By offloading heavy multimodal inference to the cloud, the system mitigates computational congestion, supporting operational resilience despite the inevitable physical bandwidth limitations of field deployments.

Zhong-Hao Zhou, Hai-Lu Xin, Ping Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.