Sep 2026· Italian National Conference on Sensors· Vol 26· 0 citations· 71 references
Medicine
TL;DR
DiMMPose is proposed, a diffusion-based framework enhanced by Mamba’s state-space model for robust and efficient 3D pose estimation, which combines structured prompts encoded by LongCLIP with learnable prompt representations to provide anatomical and motion-related guidance during denoising.
Abstract
Monocular 3D Human Pose Estimation (3D HPE) typically adopts a two-stage approach: estimating 2D joint positions from images and then lifting them to 3D coordinates, effectively reducing dataset bias inherent in direct methods. However, current lifting techniques face two key challenges: many Transformer-based methods rely on attention-based or staged spatial–temporal modeling, which can limit efficient long-range frame-joint reasoning, while diffusion models support probabilistic modeling of pose uncertainty but remain sensitive to joint-coordinate noise. We propose DiMMPose, a diffusion-based framework enhanced by Mamba’s state-space model for robust and efficient 3D pose estimation. Its denoising process consists of two coordinated modules. The Spatiotemporal Mamba Block (STMB) serves as the core feature extraction module, employing internal Pose Mamba components with bidirectional state propagation and linear complexity to efficiently model long-range frame-joint dependencies. STMB further refines these features through Spatiotemporal Scan and Merge, which traverses the same skeleton tokens in complementary frame-joint orders and fuses the resulting representations. The Multi-Prompt Mamba Denoiser (MPMD) combines structured prompts encoded by LongCLIP with learnable prompt representations to provide anatomical and motion-related guidance during denoising. DiMMPose achieves an average MPJPE of 28.9 mm on Human3.6M under the DET setting, with action-specific errors of 21.2 mm for Walking and 22.0 mm for WalkTogether. It improves over FinePOSE by 3.0 mm, reduces inference latency by 57.1%, and achieves 23.0 mm MPJPE on MPI-INF-3DHP (N = 243).
Transformers have become dominant in 3D Human Pose Estimation (HPE). However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work we leverage recent advances in state space models and utilize Mamba for high-qua...
Xin-Yi Zhang, Qi-Qi Bao, Wen-Ming Yang et al.· ACM Transactions on Multimed...· 0 citations
While single-view 3D reconstruction has seen significant progress, extrapolating complex 3D structures from inherently ambiguous 2D observations remains fundamentally ill-posed, particularly in the critically underexplored data-scarce regime. To address this challenge, we propose Point Diffusion Mamba (PDM), a method t...
Wei Zhou, Xin-Zhe Shi, Xingxing Hao et al.· 0 citations
This paper proposes a lightweight high-resolution pose estimation network, UG-HRNet, which significantly reduces model complexity while maintaining competitive performance, serving as a practical alternative to existing popular lightweight networks.
Yong-Feng Qi, Jun-Teng Zhang· International Journal of Mac...· 0 citations
Monocular 3D Human Pose Estimation (3D HPE) is a core task in computer vision, which provides critical technical support for applications including VR/AR, immersive human-computer interaction, digital human animation, as well as sports motion capture, athletic technique diagnosis, physical fitness assessment and sports...
CMambaDepth is proposed, a self-supervised framework that achieves efficient multi-scale feature fusion and fine-grained contextual modeling via channel-wise selective state propagation and a Hybrid Attention Module is introduced to combine large-kernel local context and Manhattan self-attention for complementary spati...
Xue-Lian Xiang, Jia-Yao Liu, He-Qi Xiang et al.· 0 citations
B2TFPose is presented, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.
Ali Rafiaei, Michael A. Greenspan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.