Skip to content
Conference

EVLM: Intent-Driven Edge Vision Language Model for UAV-Based Power Line Inspection

Jul 2026 · International Conference on Edge Computing [Services Society] · pp. 32-42 · 1 citation · 46 references

Abstract

Inspection of critical infrastructure, such as power lines, is increasingly conducted using unmanned aerial vehicles (UAVs) that capture aerial video for subsequent human review. Although recent edge-based approaches deploy onboard object detectors to identify predefined defect classes, these pipelines remain closed-set, task-specific, and largely decoupled from operator intent and edge resource constraints. This paper introduces EVLM, an intent-driven vision-language framework for onboard UAV-based power line inspection. Given a high-level operator intent, EVLM (i) leverages lightweight histogram-based frame filtering to extract salient key frames under bounded compute budgets, (ii) executes a domain-adapted vision language model (VLM) directly on the UAV for intent-conditioned multimodal reasoning, and (iii) synthesizes structured inspection reports together with a minimal set of evidence frames, replacing continuous raw video transmission with compact semantic outputs. To align the VLM with infrastructure inspection semantics while preserving edge efficiency, we perform parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA), enabling domain specialization without updating the full model parameters. We implement and fully deploy EVLM on an NVIDIA Jetson device representative of UAV-class onboard hardware and evaluate it using 20 publicly released power line inspection video sequences spanning 8 heterogeneous environments and 5 operational intent categories. Experimental results show a data reduction of 94.8%, with transmitted data decreasing from 485kB to 25kB per 4s segment, corresponding to 72.75MB versus 3.75MB over a 10min inspection mission. EVLM operates feasibly on embedded hardware, maintaining moderate CPU/GPU utilization and bounded power consumption (5.6W), while producing interpretable, intent-aligned inspection outputs. with richer semantic insights than detection-centric baselines.

View source

Similar papers

Preprint Sep 2026

From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection

Critical infrastructure assets such as bridges, tunnels, dams, and power line networks require timely and scalable inspection. While conventional manual inspection remains costly and hazardous, unmanned aerial vehicle (UAV)-based inspection has emerged as an efficient alternative for monitoring difficult-to-access stru...

Reza Farahani, Naser Hossein Motlagh, Zoha Azimi et al. · 0 citations
Review Open access 2026

Edge-Ready Multimodal and Geometry-Aware Perception for UAV-Based High-Voltage Power Line Inspection: A Systematic Review

Unmanned Aerial Vehicles (UAVs) have become a promising alternative for high-voltage power line inspection because they reduce operational risk, inspection time, and human exposure to hazardous environments. However, reliable real-time power line perception remains a critical bottleneck for autonomous deployment, parti...

Sebastian Aucapina, Viviana Moya, William Chamorro et al. · 0 citations
Review Aug 2026

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

MoRAL (Multimodal Reasoning for Autonomous Language Models), a two-stage fine-tuning pipeline that teaches Cosmos-Reason2-2B to first read a physics-encoded Bird's Eye View (BEV) representation and then reason over it for driving decisions, establishes a reproducible foundation for compact, physics-grounded VLM reasoni...

Ambarish Govindarajulu Kaliamurthi, Kai Liu · 0 citations
Preprint Aug 2026

TriCLE: Tri-Modal Vision-Language Reasoning for Edge-Deployed Fine-Grained Clustering

The results support TriCLE as a practical prototype for interpretable, edge-feasible aircraft grouping, while emphasizing the need for further validation on real aligned thermal and LiDAR sensor streams.

K. Gupta, Md. Mahfuzur Rahman, Fahad Rahman et al. · 0 citations
Preprint Aug 2026

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

UAV-MAS is proposed, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error...

Hao-Yu Zhang, Shuoxun Zhang, Peng Ye et al. · 0 citations
#edge computing Review Open access Sep 2026

Large Language Model-Driven Autonomous UAV Systems: Technical Evolution, Core Architectures, and Critical Challenges

The rapid development of large language models (LLMs) has expanded the capabilities of autonomous unmanned aerial vehicle (UAV) systems in naturallanguage instruction understanding, multimodal perception, and decision-making. This survey reviews the technical evolution, system architectures, and deployment challenges o...

Mei-Jie Zhang, Hao Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.