Jul 2026· Engineering Research Express· Vol 8, pp. 155213· 0 citations· 30 references
Physics
TL;DR
SMG-YOLO is presented, an enhanced YOLO-based pose estimation framework built upon Hyper-YOLO-Pose that provides a favorable accuracy–complexity trade-off for human pose estimation.
Abstract
Human pose estimation is a fundamental task in computer vision that aims to localize key human joints in images. Although you only look once (YOLO)-based pose estimation methods provide advantages in computational efficiency and deployment convenience, they still face challenges in complex backgrounds, occlusions, small-scale keypoints, and structural inconsistency among predicted joints. To address these issues, this paper presents SMG-YOLO, an enhanced YOLO-based pose estimation framework built upon Hyper-YOLO-Pose. The proposed model integrates three pose-oriented components: the selective boundary aggregation (SBA) module, the mixed aggregation network (MANet)-StarC module, and a GroupNorm-based pose detection head, namely group normalization (GN)-Pose. The SBA module is adopted to strengthen semantic-spatial feature interaction between high-level semantic features and low-level spatial cues. The MANet-StarC module incorporates the StarC unit into the MANet structure and combines nonlinear feature modulation with context anchor attention to enhance contextual feature representation. GN-Pose introduces GN into the pose prediction head to improve keypoint regression stability. Experiments on the MPII human pose dataset show that SMG-YOLO achieves 85.30% AP50 and 47.70% AP50:95, improving the Hyper-YOLO-Pose baseline by 2.30 and 2.40 percentage points, respectively. Additional ablation experiments further verify the complementary effects of the proposed components. Moreover, model-forward speed testing on an RTX 3060 GPU shows that SMG-YOLO achieves 49.1 FPS, indicating practical inference efficiency under the tested setting. These results demonstrate that SMG-YOLO provides a favorable accuracy–complexity trade-off for human pose estimation.
Human pose estimation in factory surveillance is challenged by scale variation, limb occlusion, complex backgrounds, and spatial detail loss during upsampling. This paper proposes MFA-Pose, an improved YOLO11s-Pose framework that enhances contextual representation and cross-scale feature reconstruction. The Multi-Scale...
This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.
Xinmiao Du· Poster Volume 0007 The 2026...· 0 citations
Pose estimation is a key aspect of action recognition, study of behaviors and spatial relationships from video data. Traditional pose estimation methods are designed exclusively for keypoint detection of either humans or vehicles and incur high inference time and complexity. In this paper, we present UnifiedPoseNet, a...
Reenie Tanya, Balika J. Chelliah· IEEE Access· 0 citations
Human pose estimation is an important task in computer vision and has been widely applied in many fields. Existing algorithms have achieved promising performance in regular motion scenarios, but estimating poses under irregular motions caused by body flipping, folding, fast motion, and occlusion remains challenging. To...
Jiahong Jiang, Nan Xia, Mengyuan Wang· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.