Skip to content
Open access

A Lightweight GDMM-YOLO11 Model for Cotton–Weed Instance Segmentation and Image-Plane Operation-Point Localization

Sep 2026 · Agronomy · 0 citations · 31 references

Abstract

To address the challenges of crop–weed instance segmentation and image-plane operation-point localization under visual similarity, background interference, leaf occlusion, and irregular plant morphology in cotton fields, a lightweight instance-segmentation model, GDMM-YOLO11, was developed. Based on YOLO11n-seg, the four backbone stage-transition downsampling convolutions at P2/4, P3/8, P4/16, and P5/32 were replaced with GhostConv to reduce redundant computation, a C3k2_DySnakeConv_Mona module was introduced to strengthen structural and multi-scale feature representation, and the original post-SPPF C2PSA block was replaced with mixed local channel attention (MLCA) to recalibrate high-level feature responses. Experiments were conducted on 2177 images containing 2856 annotated plant instances using a stratified 65%/15%/20% training–validation–test split. Across three independent runs with random seeds 3407, 3408, and 3409, GDMM-YOLO11 achieved a mask precision of 89.82 ± 2.60%, mask recall of 84.50 ± 2.49%, mask mAP@0.5 of 89.76 ± 0.66%, and mask mAP@0.5:0.95 of 68.50 ± 0.34%. Relative to YOLO11n-seg, the corresponding three-run mean values were numerically higher by 3.52, 1.07, 1.67, and 2.33 percentage points, respectively, while the parameter count decreased from 2.836 M to 2.533 M and the computational cost decreased from 9.6 to 8.8 GFLOPs. For image-plane operation-point localization, 505 of 528 ground-truth weed instances obtained valid same-class mask matches. The proposed skeleton-constrained fused-center method achieved a ground-truth-mask inclusion rate of 99.41%, a conditional Success@0.15 of 91.29%, and an end-to-end Success@0.15 of 87.31%. Paired comparisons with the mask-centroid baseline showed statistically significant improvements in ground-truth-mask inclusion and weed-boundary clearance, whereas differences in localization error and Success@0.10/0.15 were not statistically significant. TensorRT FP16 deployment on an NVIDIA Jetson AGX Orin achieved a model-only inference latency of 2.796 ± 0.115 ms, corresponding to 357.67 FPS. These results show that GDMM-YOLO11 provides a favorable accuracy–complexity trade-off while supporting image-plane operation-point generation and high-throughput model-only edge inference.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.