Results support acquisition-aware same-test evaluation as a necessary complement to ordinary image-level splitting in multimodal UAV benchmarks and reveal strong near-sequential dependence.
Abstract
UAV image collections contain spatially and temporally related frames, yet semantic-segmentation benchmarks commonly split them at image level. Such splitting can place samples from one acquisition in both model development and testing, obscuring transfer to a genuinely new survey. Using the 734-sample WeedyRice-RGBMS-DB, we fix a 124-image target-acquisition test set and compare two protocols with identical train, validation, and test counts: target-held-out, which excludes the target acquisition from development, and target-exposed, which admits its remaining images. SegFormer-B0 is evaluated with RGB, four-band multispectral (MS), and seven-channel RGB+MS input over two fixed-split seeds. RGB is strongest under complete acquisition holdout ($0.7317\pm0.0201$ IoU), whereas RGB+MS becomes strongest after target exposure ($0.7822\pm0.0269$). A fixed-split U-Net/ResNet18 replication confirms positive exposure gains for all three inputs, but retains RGB as the best modality under both protocols. Acquisition exposure therefore increases measured performance across both evaluated backbones, while its effect on modality ranking is architecture-dependent. A supplied-split audit reveals strong near-sequential dependence, and corruption tests show that early fusion is substantially more sensitive to RGB--MS displacement than to moderate radiometric scaling. These results support acquisition-aware same-test evaluation as a necessary complement to ordinary image-level splitting in multimodal UAV benchmarks. The code and supporting the findings of this study will be publicly released upon acceptance of the paper.
Unmanned Aerial Vehicles (UAVs) have become widely used in agricultural monitoring and precision agriculture. In real field conditions, however, UAV-based analysis is often affected by dense planting patterns, background interference, and unclear boundaries, which make reliable computer vision analysis more difficult....
Ze-Yu Ren· International Conference on...· 0 citations
Unmanned Aerial Vehicles (UAVs) have emerged as a promising platform for firefighting operations due to their flexibility, low operational cost, and ability to acquire high-resolution imagery in locations that may be difficult or dangerous to access using conventional methods. Recent advances in deep learning have sign...
Matheus F. Kovaleski, L. Garrote, C. Premebida et al.· 0 citations
Crop and forest damage from sika deer, wild boar, and Japanese macaque is a serious economic problem in Japanese satoyama, where farmland and forest intermingle. Camera traps enable scalable monitoring, yet existing benchmarks evaluate recognition in-distribution, rarely prioritize night infrared imagery, and do not jo...
Keito Inoshita, Kohei Hisayama, Haruto Sugeno et al.· 0 citations
Traditional feature-based image stitching methods depend heavily on the quality of feature matching, which leads to suboptimal performance when applied to drone remote sensing images with significant differences in viewpoint and depth of field. Concurrently, supervised learning paradigms have proven infeasible due to t...
Wenpeng Zhang, Xiangyue Zhang, Huaguang Shi et al.· IEEE Geoscience and Remote S...· 0 citations
This application study evaluates two existing instance-segmentation frameworks, RF-DETR (Region Focused Detection Transformer) and YOLOv26, for rip current detection in complex marine imagery. Both models were fine-tuned from pretrained weights on the same 18,389 training images. YOLOv26 used the 4349 labeled images as...
Van Lam Ho, Van Khang Le, Xuan Vinh Le et al.· Artificial Intelligence and...· 0 citations
Accurate pavement-distress detection in unmanned aerial vehicle (UAV) imagery remains challenging because defects exhibit weak texture, irregular geometry, scale variation, and visual similarity to markings, shadows, water, and repaired pavement. This study presents YOLO-MCG, an engineering adaptation of YOLOv10 that i...
Rong Li, Jian Liu, Cui-Zhen Sun et al.· Electronics· 0 citations
Related blog posts
Microsoft Research Blog· microsoft.comAug 11, 2026
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.