Jul 2026· International Conference on Image Processing and Intelligent Control· Vol 14262, pp. 142620K - 142620K-7· 0 citations· 11 references
Engineering
TL;DR
This study presents an enhanced YOLOv11 framework specifically designed for global geometric perception and high-fidelity single-stage 6D pose regression, which validate that the integration of global geometric awareness consistently outperforms the vanilla YOLOv11 and other classical baselines in complex scenarios.
Abstract
Precise 6D object pose estimation from RGB images remains a formidable challenge due to complex backgrounds and severe occlusions. To address these issues, our study presents an enhanced YOLOv11 framework specifically designed for global geometric perception and high-fidelity single-stage 6D pose regression. The core of our architecture is the C3k2SW module, which innovatively synergizes local convolutional features with global long-range dependencies through windowbased self-attention, significantly enhancing the network's geometric perception of spatial topologies. Furthermore, to optimize multi-scale feature interaction, an adaptive ConcatA module and a Bi-directional Feature Pyramid Attention Network (BFPAN) are proposed to suppress background noise while preserving fine-grained geometric details across different scales. Experimental results on the LineMod benchmark demonstrate that our method achieves an optimal tradeoff between inference efficiency and accuracy, reaching an average ADD(-S) accuracy of 76.50% and 84.92% on the 5cm 5° metric, respectively. These results validate that the integration of global geometric awareness consistently outperforms the vanilla YOLOv11 and other classical baselines in complex scenarios.
This study introduces a novel optimization framework for YOLOv11, specifically engineered for tiny-scale targets by integrating convolutional block attention modules (CBAM), k-means anchor clustering, and an enhanced feature pyramid network (FPN).
H. Husin, H. Hao, Pan-Yuan Fei et al.· IAES International Journal o...· 0 citations
Conditional flow matching has enabled a step forward in object 6D pose estimation, achieving state-of-the-art performance by progressively denoising and registering object representations to observed scenes. Existing methods require training task-specific encoders supervised on object-scene overlap and rely on trivial...
Amir Hamza, Davide Boscaini, Fabio Poiesi· 0 citations
A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.
Jie Li, Jia-Heng Xu, Laiyan Ding et al.· International Conference on...· 0 citations
This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.
Xinmiao Du· Poster Volume 0007 The 2026...· 0 citations
Achieving accurate and efficient object pose estimation is a key goal in computer vision. Most existing methods rely on controlled environments, limiting their effectiveness in complex, dynamic, and unstructured real-world scenarios, especially for novel objects, severe occlusion, or sensor noise. Recent studies show t...
Hui Zhang, Yue Wang, Jianhao Jiao et al.· IEEE Transactions on Automat...· 0 citations
High-precision perception is fundamental to safe autonomous driving, and BEV-based 3D object detection via lidar-camera fusion plays a crucial role in improving detection accuracy and robustness. To address insufficient feature representation, spatial misalignment, and the limitations of static fusion strategies, this...
Jie Hu, Xinghao Cheng, Shuaidi He et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.