Jul 2026· 2026 6th International Conference on Intelligent Communications and Computing (ICICC)· pp. 1-6· 0 citations· 16 references
TL;DR
These findings confirm that context-aware adaptive scheduling can dramatically boost the energy efficiency of edge multimodal inference without sacrificing task performance, offering a viable deployment solution for low-power Internet of Things, intelligent surveillance, and human-robot interaction scenarios.
Abstract
Deploying multimodal artificial intelligence models on resource-constrained edge devices faces inherent bottlenecks in energy consumption and computational latency, as conventional full-modality inference pipelines keep all perception encoders active regardless of task context and environmental conditions, causing substantial power waste and degraded real-time performance. To address this challenge, this work presents an energy-efficient context-aware multimodal edge inference framework featuring a lightweight modality activation sparsity evaluation unit, dynamic computation path scheduling, and a cross-modal speculative skipping mechanism. The framework quantifies the information value of each input modality in real time according to scene context, task complexity, and device power status, and adaptively activates or deactivates corresponding visual, audio, and sensor encoders, while tuning model quantization precision and operator fusion strategies to align with runtime resource budgets. Validated on NVIDIA Jetson Nano and Raspberry Pi 5 edge platforms across VQAv2, MMBench, and multimodal perception benchmarks, the proposed framework delivers a 42.3% reduction in end-to-end energy consumption and a 30%-65% decrease in inference latency against static full-modality baselines, alongside $\mathbf{1. 5} \times$ to $\mathbf{2. 3} \times$ higher throughput with an accuracy loss no more than 1.2%. The runtime context scheduling module introduces less than 9 ms of latency overhead and only 0.32 W of additional power draw, with per-inference energy as low as 0.6 J; for battery-powered mobile edge devices, the framework extends continuous operating duration by over 72% under typical perception workloads. These findings confirm that context-aware adaptive scheduling can dramatically boost the energy efficiency of edge multimodal inference without sacrificing task performance, offering a viable deployment solution for low-power Internet of Things, intelligent surveillance, and human-robot interaction scenarios.
A compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference and demonstrates the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perceptio...
Riadul Islam, Joey Mulé, Dhandeep Challagundla et al.· 0 citations
Neuro-Elastic, an adaptive inference framework that operates at two granularities: per-input sparsity (entropy-driven early exit and token pruning) and device-state-driven switching among pre-compiled mixed-precision model variants, is proposed.
Wenbin Shang, Dai Teng, Tingjie Chen et al.· IEEE Access· 1 citation
This work presents an energy-proportional, context-aware vision IoT node that addresses this challenge through a heterogeneous multimodal dual-camera architecture, enabling always-on visual monitoring in a place-and-forget scenario through autonomous edge intelligence.
Julian Moosmann, P. Mayer, Luca Benini et al.· 0 citations
Recent advances in multimodal large language mod- els (MLLMs) have opened new opportunities for edge intelligence by enabling reasoning across heterogeneous sensor modalities, such as vision, text, and telemetry data. However, deploying these capabilities on resource-constrained edge platforms remains challenging due t...
Multimodal AI streaming is used to enable real-time interaction in educational applications. It is a critical component in the integration of AI-driven Augmented Reality (AR) for language learning. The challenge faced in AI streaming is maintaining responsiveness on resource-constrained devices in regions with unstable...
A. Rahman· International Journal of Ele...· 0 citations
On-device model adaptation is essential to enable lifelong personalization on resource-constrained hardware, but compute, power, and memory limitations of such devices make end-to-end backpropagation impractical for modern deep neural networks. This work proposes a heterogeneous adaptation pipeline that repurposes a co...
M. Piechocki, Alessandro Capotondi, Marek Kraft· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.