This work presents a monocular 6D pose estimation approach using hybrid representations and adaptive multi-feature fusion to address challenges of augmented reality-assisted assembly, human–robot collaboration, and quality inspection in intelligent manufacturing.
Abstract
Vision-based 6D pose estimation is critical for augmented reality-assisted assembly, human–robot collaboration, and quality inspection in intelligent manufacturing. However, performance degrades severely in complex industrial scenarios due to occlusion, varying lighting, textureless surfaces, and reflective parts. This work presents a monocular 6D pose estimation approach using hybrid representations and adaptive multi-feature fusion to address these challenges. A hybrid representation learning framework is designed to jointly predict keypoint heatmaps, relational vectors, semantic edges, masks, and visibility, thereby enhancing feature robustness. A multi-feature adaptive fusion strategy optimizes the pose by combining semantic and fine-grained general features. A structure-constrained correction module refines multi-object poses using assembly consistency constraints. Experiments on a custom industrial assembly dataset and the public Mono6D dataset show that the proposed method achieves 87.46% ADD (0.1d) and 86.82% 5 cm/5∘\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^\circ $$\end{document} accuracy, outperforming state-of-the-art methods. The custom dataset includes multiple weakly textured and reflective assembly parts under occlusion, lighting variation, and multi-viewpoint conditions. Furthermore, the complete system runs at approximately 18 FPS, with faster tracking once initialized. The approach supports reliable AR-assisted assembly and meets industrial deployment requirements. Our code and datasets are open-sourced at https://github.com/nengbinlv/HRMFPose, with the DOI: https://doi.org/https://doi.org/10.5281/zenodo.19574143.
A two-stage method for accurate 6D pose estimation using consumer-grade depth cameras in real-world scenarios with YOLO-based detection, Euclidean clustering, and moving least squares smoothing combined to extract high-quality target point clouds from noisy RGB-D observations is proposed.
This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.
Xinmiao Du· Poster Volume 0007 The 2026...· 0 citations
Relative pose estimation under the query–reference paradigm has emerged as a practical alternative for estimating the 6D pose of unseen objects without relying on CAD models or extensive annotations. However, many existing approaches remain difficult to deploy in practice, as their reliance on geometric matching make...
Jia-Le Ren, Ming-Xing Tan, Meng-Yuan Liu et al.· CAAI Transactions on Intelli...· 0 citations
A depth-guided lightweight multi-task 3D scene understanding method based on a lightweight U-Net architecture, uses RGB and depth four-channel dual- modal input, and constructs a network structure with a shared encoder and independent decoders, allowing a single inference to simultaneously accomplish the three core tas...
A real-time point cloud processing and workpiece localization system that integrates multi-module optimization with a dynamic adaptive framework that sustains reliable performance under Gaussian noise up to 0.6 mm and occlusion levels up to 40%, confirming its viability for high-throughput industrial applications.
3D object detection is an essential and highly challenging task for unmanned autonomous systems operating in indoor scenes. Current mainstream 3D detection approaches rely on the direct encoding of point cloud coordinates, while neglecting the underlying geometric priors, thereby limiting the expressiveness of learne...
Jian-Jun Ni, Sheng-Ying Wu, Jie Liu et al.· Measurement science and tech...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.