Skip to content
Open access

Hybrid representation and adaptive multi-feature fusion for monocular 6D pose estimation in industrial assembly

Sep 2026 · The Visual Computer · Vol 42 · 0 citations · 80 references

TL;DR

This work presents a monocular 6D pose estimation approach using hybrid representations and adaptive multi-feature fusion to address challenges of augmented reality-assisted assembly, human–robot collaboration, and quality inspection in intelligent manufacturing.

Abstract

Vision-based 6D pose estimation is critical for augmented reality-assisted assembly, human–robot collaboration, and quality inspection in intelligent manufacturing. However, performance degrades severely in complex industrial scenarios due to occlusion, varying lighting, textureless surfaces, and reflective parts. This work presents a monocular 6D pose estimation approach using hybrid representations and adaptive multi-feature fusion to address these challenges. A hybrid representation learning framework is designed to jointly predict keypoint heatmaps, relational vectors, semantic edges, masks, and visibility, thereby enhancing feature robustness. A multi-feature adaptive fusion strategy optimizes the pose by combining semantic and fine-grained general features. A structure-constrained correction module refines multi-object poses using assembly consistency constraints. Experiments on a custom industrial assembly dataset and the public Mono6D dataset show that the proposed method achieves 87.46% ADD (0.1d) and 86.82% 5 cm/5∘\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$^\circ $$\end{document} accuracy, outperforming state-of-the-art methods. The custom dataset includes multiple weakly textured and reflective assembly parts under occlusion, lighting variation, and multi-viewpoint conditions. Furthermore, the complete system runs at approximately 18 FPS, with faster tracking once initialized. The approach supports reliable AR-assisted assembly and meets industrial deployment requirements. Our code and datasets are open-sourced at https://github.com/nengbinlv/HRMFPose, with the DOI: https://doi.org/https://doi.org/10.5281/zenodo.19574143.

Read PDF

Similar papers

Aug 2026

Accurate 6D pose estimation using consumer-grade depth cameras in real-world scenarios

A two-stage method for accurate 6D pose estimation using consumer-grade depth cameras in real-world scenarios with YOLO-based detection, Euclidean clustering, and moving least squares smoothing combined to extract high-quality target point clouds from noisy RGB-D observations is proposed.

Zhi-Xiang Zhang, Feng Luan, Wen-Hao Li · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.

Xinmiao Du · 0 citations
Open access Sep 2026

MRPose: Multi‐Robot Relative 6D Pose Estimation for Unseen Objects From RGB Images

Relative pose estimation under the query–reference paradigm has emerged as a practical alternative for estimating the 6D pose of unseen objects without relying on CAD models or extensive annotations. However, many existing approaches remain difficult to deploy in practice, as their reliance on geometric matching make...

Jia-Le Ren, Ming-Xing Tan, Meng-Yuan Liu et al. · 0 citations
Conference Open access 2026

Deeply Guided Lightweight 3D Scene Understanding Method and Its Application in Mobile Robot Perception

A depth-guided lightweight multi-task 3D scene understanding method based on a lightweight U-Net architecture, uses RGB and depth four-channel dual- modal input, and constructs a network structure with a shared encoder and independent decoders, allowing a single inference to simultaneously accomplish the three core tas...

Chengkai Shi · 0 citations
Open access Aug 2026

An integrated 3D vision sensor system with adaptive point cloud processing for robotic workpiece localization in industrial bin picking

A real-time point cloud processing and workpiece localization system that integrates multi-module optimization with a dynamic adaptive framework that sustains reliable performance under Gaussian noise up to 0.6 mm and occlusion levels up to 40%, confirming its viability for high-throughput industrial applications.

Zhi-Wen Xiong, Yuanchun Li · 0 citations
Open access Sep 2026

An improved geometric priors-based indoor 3D object detection in point clouds for mobile robots

3D object detection is an essential and highly challenging task for unmanned autonomous systems operating in indoor scenes. Current mainstream 3D detection approaches rely on the direct encoding of point cloud coordinates, while neglecting the underlying geometric priors, thereby limiting the expressiveness of learne...

Jian-Jun Ni, Sheng-Ying Wu, Jie Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.