Skip to content
Conference Open access

Energy-Efficient Context-Aware Multimodal AI Inference at the Edge

Jul 2026 · 2026 6th International Conference on Intelligent Communications and Computing (ICICC) · pp. 1-6 · 0 citations · 16 references

TL;DR

These findings confirm that context-aware adaptive scheduling can dramatically boost the energy efficiency of edge multimodal inference without sacrificing task performance, offering a viable deployment solution for low-power Internet of Things, intelligent surveillance, and human-robot interaction scenarios.

Abstract

Deploying multimodal artificial intelligence models on resource-constrained edge devices faces inherent bottlenecks in energy consumption and computational latency, as conventional full-modality inference pipelines keep all perception encoders active regardless of task context and environmental conditions, causing substantial power waste and degraded real-time performance. To address this challenge, this work presents an energy-efficient context-aware multimodal edge inference framework featuring a lightweight modality activation sparsity evaluation unit, dynamic computation path scheduling, and a cross-modal speculative skipping mechanism. The framework quantifies the information value of each input modality in real time according to scene context, task complexity, and device power status, and adaptively activates or deactivates corresponding visual, audio, and sensor encoders, while tuning model quantization precision and operator fusion strategies to align with runtime resource budgets. Validated on NVIDIA Jetson Nano and Raspberry Pi 5 edge platforms across VQAv2, MMBench, and multimodal perception benchmarks, the proposed framework delivers a 42.3% reduction in end-to-end energy consumption and a 30%-65% decrease in inference latency against static full-modality baselines, alongside $\mathbf{1. 5} \times$ to $\mathbf{2. 3} \times$ higher throughput with an accuracy loss no more than 1.2%. The runtime context scheduling module introduces less than 9 ms of latency overhead and only 0.32 W of additional power draw, with per-inference energy as low as 0.6 J; for battery-powered mobile edge devices, the framework extends continuous operating duration by over 72% under typical perception workloads. These findings confirm that context-aware adaptive scheduling can dramatically boost the energy efficiency of edge multimodal inference without sacrificing task performance, offering a viable deployment solution for low-power Internet of Things, intelligent surveillance, and human-robot interaction scenarios.

Read PDF

Similar papers

#edge computing Preprint Aug 2026

LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

A compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference and demonstrates the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perceptio...

Riadul Islam, Joey Mulé, Dhandeep Challagundla et al. · 0 citations
Open access 2026

Neuro-Elastic: A Unified Framework for Hardware-Aware Adaptive Quantization and Dynamic Sparsity in Real-Time Edge Intent Prediction

Neuro-Elastic, an adaptive inference framework that operates at two granularities: per-input sparsity (entropy-driven early exit and token pruning) and device-state-driven switching among pre-compiled mixed-precision model variants, is proposed.

Wenbin Shang, Dai Teng, Tingjie Chen et al. · 1 citation
Preprint Aug 2026

An Energy-Proportional Multimodal and Context-Aware Vision IoT Node

This work presents an energy-proportional, context-aware vision IoT node that addresses this challenge through a heterogeneous multimodal dual-camera architecture, enabling always-on visual monitoring in a place-and-forget scenario through autonomous edge intelligence.

Julian Moosmann, P. Mayer, Luca Benini et al. · 0 citations
#machine learning Preprint Sep 2026

EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression

Recent advances in multimodal large language mod- els (MLLMs) have opened new opportunities for edge intelligence by enabling reasoning across heterogeneous sensor modalities, such as vision, text, and telemetry data. However, deploying these capabilities on resource-constrained edge platforms remains challenging due t...

Motahare Mounesan, Irfan Khan · 0 citations
Open access Jul 2026

Real-time edge intelligence using State-Gated Spectral VAD for multimodal streaming on resource-constrained devices

Multimodal AI streaming is used to enable real-time interaction in educational applications. It is a critical component in the integration of AI-driven Augmented Reality (AR) for language learning. The challenge faced in AI streaming is maintaining responsiveness on resource-constrained devices in regions with unstable...

A. Rahman · 0 citations
Jul 2026

Empowering On-Device Model Adaptation with an Edge AI Inference Accelerator

On-device model adaptation is essential to enable lifelong personalization on resource-constrained hardware, but compute, power, and memory limitations of such devices make end-to-end backpropagation impractical for modern deep neural networks. This work proposes a heterogeneous adaptation pipeline that repurposes a co...

M. Piechocki, Alessandro Capotondi, Marek Kraft · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.