Improving UAV Assisted Spatial Monitoring with Semiautomated Fine Tuning of Vision Language Models
Abstract
Detecting low-metal anti-personnel mines (e.g., PFM-1) challenges traditional clearance methods, while standard UAV computer vision often generates high falsepositive rates in cluttered environments. To address this, we propose a multilevel edge-ground-cloud architecture that escalates ambiguous ground-station detections to a cloud-based Vision Language Model (VLM) semantic verifier. Through a novel human-in-the-loop mechanism, operators continuously fine-tune these models using confirmed, contextually padded image crops. Evaluations demonstrate that the fine-tuned GPT-4.1 model improves detection accuracy by 42.7% over baselines (outperforming GPT-4o by 11.6%), reducing errors to just 63 misclassifications across 339 frames. Costing approximately $3-$4 per iteration, this efficient pipeline drastically suppresses false alarms while preserving high recall, democratizing specialized AI for global humanitarian demining.