Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-7· 0 citations· 42 references
Abstract
D pose estimation in densely-packed, industrial environments remains a challenging perception problem due to severe occlusions, sensor noise, and the fragmentary nature of reconstructed 3D observations. To this end, we present a generative module for holistic, depth-only 6D pose estimation in cluttered bin-picking that integrates a conditioned Denoising Diffusion Probabilistic Model (DDPM) as a learned data-centric inference adaptor for point cloud segments. Building on a holistic staged-heatmap fusion backbone that produces focused, memoryefficient point cloud fragments around candidate objects, we show that sharpening and completing these fragments with a diffusion-based point cloud autoencoder substantially improves downstream per-point voting pose estimation. Our method integrates a velocity-targeted DDPM, implemented via a Point Transformer V3 (PTv3) backbone, to reconstruct sharper and more complete object-level geometry from imperfect scene fragments. The denoiser is trained to explicitly align geometric completion with the 6D pose objective. At inference we apply a lightweight sampling schedule to produce a single sharpened segment per object candidate, preserving time-efficiency for robotics bin-picking. With pilot experiments on the IPD dataset, we demonstrate that diffusion-based completion reduces pose error and increases robustness to occlusion, sensor noise, and severe fragmentation, while retaining the holistic advantages of scene-level reasoning. We argue that diffusion models are a promising direction for scene denoising and completion in 6D pose estimation and provide a practical integration strategy for robotic perception systems.
Map-Det3D is an online multi-view 3D object detection model that brings detection directly into a 3D space reconstructed from RGB, suggesting that training reconstruction priors for detection is a practical route to stable metric 3D detection from monocular video.
Yung-Hsu Yang, Luigi Piccinelli, S. R. Bulò et al.· 1 citation
PIXIE is a zero-shot framework that estimates the 6D pose of an object from an RGB image using only an untextured 3D model, inherently robust to lighting and texture variation, while correspondence filtering handles geometric deviations between the model and physical object.
Leon Jungemeyer, A. Magaña, Gautham Mohan et al.· arXiv.org· 0 citations
WaveletkAN is a pose-refinement model based on RGB-D point clouds proposed in this paper to correct object pose hypotheses for dense bins with occlusion, specular depth noise, and self-similar industrial parts and can achieve robust and deployable pose correction for RGB-D robotic bin-picking systems.
Charalampos Evangelou, Iakovos Maniatis· Journal of Applied Automatio...· 0 citations
DOU-Pose is proposed, a visual pose estimation framework built upon the Differentiable SAmple Consensus (DSAC)* pipeline to enhance the discriminative capability of scene coordinate regression through improved feature extraction and replaces standard convolutional layers with Depthwise Over-parameterized Convolution (D...
Xin'an Qiu, Li-Wen Wang, Zezheng Dong et al.· Italian National Conference...· 0 citations
FlexSplat matches or approaches posed state-of-the-art reconstructors while requiring neither camera poses nor ground-truth depth, and matches the best perceptual (LPIPS) quality among the compared methods on GSO.
Amir Sabbaghziarani, Han-Ting Ye, Maria Gorlatova et al.· 0 citations
Image-to-Point Cloud Registration aims to estimate the camera pose of a given image within a 3D scene point cloud, which is a fundamental task in autonomous driving and large-scale outdoor localization. Recent implicit correspondence learning methods have improved registration performance by learning cross-modal alignm...
Wen-Xin Zhang, Hang Li, Zhiwei Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.