Skip to content
Open access

High-Fidelity Human Pose Transfer: A Unified Framework With Hierarchical Semantic Alignment and Gated Residual Fusion

2026 · IEEE Access · Vol 14, pp. 103882-103895 · 0 citations · 50 references
Computer Science

TL;DR

A unified generative adversarial network (GAN) framework that integrates three novel, complementary mechanisms that generates high-fidelity, structurally faithful results from a single reference image, offering a robust solution for applications in virtual try-on, animation, and human image synthesis under challenging pose transformations.

Abstract

Pose transfer, a core task in human-centric image generation, aims to synthesise photorealistic images of a subject in novel poses while preserving identity and intricate clothing details. Existing methods, particularly under large pose variations such as extreme articulation or self-occlusion, often struggle with preserving fine-grained textures and maintaining structural consistency, leading to artifacts like distorted limbs and lost details. To address these challenges, we introduce a unified generative adversarial network (GAN) framework that integrates three novel, complementary mechanisms. First, a Hierarchical Semantic Aligner (HSA) establishes multi-scale semantic correspondence between source appearance and target pose features through local attention and global gating. Second, a Pose-Aware Feature Injection (PAFI) module explicitly models source-target pose discrepancy to generate dynamic modulation parameters for adaptive feature adjustment during decoding. Third, a Gated Residual Fusion (GRF) strategy adaptively balances local detail and global structural information via a learnable dual-branch gating mechanism. Evaluated on the DeepFashion dataset, our framework demonstrates significant improvements, achieving a 13.9% reduction in Fréchet Inception Distance (FID) compared to the MAGPT method, alongside superior scores in Structural Similarity Index (SSIM) and Learned Perceptual Image Patch Similarity (LPIPS). Ablation studies confirm the individual contributions of each component. The proposed approach generates high-fidelity, structurally faithful results from a single reference image, offering a robust solution for applications in virtual try-on, animation, and human image synthesis under challenging pose transformations.

Read PDF

Similar papers

Jul 2026

Progressive feature-space alignment for pose-controllable virtual try-on

A feature-space alignment framework for person-to-person multi-pose virtual try-on, which estimates appearance flow in a latent feature space and uses the learned flow to warp body-part and garment-related representations to improve realism and structural fidelity.

Yang Yang, Erwei Yin · 0 citations
Open access Aug 2026

HiPA-Gen: Hierarchical Interaction Perception and Alignment for High-Fidelity Pose-Guided Human Generation

Pose-guided human image generation aims to synthesize an image of a target person based on a reference image and a target pose. Although diffusion-based methods have recently achieved significant progress in pose alignment and visual realism, generating high-fidelity images in regions involving complex pose interaction...

Juncheng Zhu, Haotian Yang, Mubai Li et al. · 0 citations
Conference Jul 2026

Face photo-sketch translation method based on global-local fast Fourier convolution and adaptive cross-domain attention

Experimental results on the CUFS and CUFSF datasets demonstrate that the proposed method achieves superior visual quality outperforms or ranks second in terms of the LPIPS, FID, and FSIM metrics, validating its effectiveness in high-quality cross-domain image generation.

Lei Zhang, Houpan Zhou · 0 citations
Jul 2026

Multi-condition guided diffusion model for face sketch-to-photo synthesis.

A diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity and compares with recent image-to-image translation and diffusion-based baselines to observe competitive performance in both visual coherence and identity preservation.

Yue Que, Xuegui Cheng, Shuqian Shi et al. · 0 citations

D 3 F: Diffusion-Driven Dual-Stream Framework for Occluded Person Re-Identification

A novel end-to-end Diffusion-Driven Dual-stream Framework (D 3 F), which seamlessly integrates generative structural priors from Diffusion Transformers (DiT) into vision-language ReID, achieving state-of-the-art (SOTA) performance on both occluded and holistic ReID benchmark datasets.

Xiaohao Xie, Weihao Meng, Wen-Hua Jiao · 0 citations
Preprint Aug 2026

UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

UniVVT is presented, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference and validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.

Yu-She Cao, Shikun Feng, Fei Shen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.