Skip to content
Open access

DiMMPose: A Diffusion-Mamba Hybrid Framework with Multi-Prompt for Efficient and Robust 3D Human Pose Estimation

Sep 2026 · Italian National Conference on Sensors · Vol 26 · 0 citations · 71 references
Medicine

TL;DR

DiMMPose is proposed, a diffusion-based framework enhanced by Mamba’s state-space model for robust and efficient 3D pose estimation, which combines structured prompts encoded by LongCLIP with learnable prompt representations to provide anatomical and motion-related guidance during denoising.

Abstract

Monocular 3D Human Pose Estimation (3D HPE) typically adopts a two-stage approach: estimating 2D joint positions from images and then lifting them to 3D coordinates, effectively reducing dataset bias inherent in direct methods. However, current lifting techniques face two key challenges: many Transformer-based methods rely on attention-based or staged spatial–temporal modeling, which can limit efficient long-range frame-joint reasoning, while diffusion models support probabilistic modeling of pose uncertainty but remain sensitive to joint-coordinate noise. We propose DiMMPose, a diffusion-based framework enhanced by Mamba’s state-space model for robust and efficient 3D pose estimation. Its denoising process consists of two coordinated modules. The Spatiotemporal Mamba Block (STMB) serves as the core feature extraction module, employing internal Pose Mamba components with bidirectional state propagation and linear complexity to efficiently model long-range frame-joint dependencies. STMB further refines these features through Spatiotemporal Scan and Merge, which traverses the same skeleton tokens in complementary frame-joint orders and fuses the resulting representations. The Multi-Prompt Mamba Denoiser (MPMD) combines structured prompts encoded by LongCLIP with learnable prompt representations to provide anatomical and motion-related guidance during denoising. DiMMPose achieves an average MPJPE of 28.9 mm on Human3.6M under the DET setting, with action-specific errors of 21.2 mm for Walking and 22.0 mm for WalkTogether. It improves over FinePOSE by 3.0 mm, reduces inference latency by 57.1%, and achieves 23.0 mm MPJPE on MPI-INF-3DHP (N = 243).

Read PDF

Similar papers

Open access Sep 2026

Pose Magic++: Integrating Mamba and HyperGCN for Efficient and Temporally Consistent 3D Human Pose Estimation

Transformers have become dominant in 3D Human Pose Estimation (HPE). However, existing Transformer-based 3D HPE backbones often encounter a trade-off between accuracy and computational efficiency. To resolve the above dilemma, in this work we leverage recent advances in state space models and utilize Mamba for high-qua...

Xin-Yi Zhang, Qi-Qi Bao, Wen-Ming Yang et al. · 0 citations
Preprint Sep 2026

Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity

While single-view 3D reconstruction has seen significant progress, extrapolating complex 3D structures from inherently ambiguous 2D observations remains fundamentally ill-posed, particularly in the critically underexplored data-scarce regime. To address this challenge, we propose Point Diffusion Mamba (PDM), a method t...

Wei Zhou, Xin-Zhe Shi, Xingxing Hao et al. · 0 citations
Sep 2026

UG-HRNet: a lightweight pose estimation network with unified guidance deformable convolution

This paper proposes a lightweight high-resolution pose estimation network, UG-HRNet, which significantly reduces model complexity while maintaining competitive performance, serving as a practical alternative to existing popular lightweight networks.

Yong-Feng Qi, Jun-Teng Zhang · 0 citations
Open access Sep 2026

HADD: A hierarchy-aware disentangled diffusion framework with spatio-temporal denoising for 3D human pose estimation

Monocular 3D Human Pose Estimation (3D HPE) is a core task in computer vision, which provides critical technical support for applications including VR/AR, immersive human-computer interaction, digital human animation, as well as sports motion capture, athletic technique diagnosis, physical fitness assessment and sports...

Jin-Ze Jiang, Zhong-Heng Jian, Hua-Ling Zheng et al. · 0 citations
Preprint Sep 2026

CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention

CMambaDepth is proposed, a self-supervised framework that achieves efficient multi-scale feature fusion and fine-grained contextual modeling via channel-wise selective state propagation and a Hybrid Attention Module is introduced to combine large-kernel local context and Manhattan self-attention for complementary spati...

Xue-Lian Xiang, Jia-Yao Liu, He-Qi Xiang et al. · 0 citations
Preprint Sep 2026

Back to the Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features

B2TFPose is presented, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.

Ali Rafiaei, Michael A. Greenspan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.