Skip to content
Open access

SSGLD: Spatial-Semantic Guided Latent Diffusion for Person Image Synthesis

2026 · IEEE Access · Vol 14, pp. 135456-135477 · 0 citations · 81 references

Abstract

Existing latent diffusion models (LDMs) for pose-guided person image synthesis (PGPIS) face an inherent trade-off between semantic consistency and texture fidelity. This dilemma stems from their reliance on a single conditioning pathway: semantic guidance is robust but prone to over-smoothing high-frequency details, whereas spatial guidance preserves textures but heavily relies on dense surface priors and lacks hallucination capabilities in occluded regions. To resolve this, we propose Spatial-Semantic Guided Latent Diffusion (SSGLD), centered on the principle of condition routing. By decoupling the synthesis task into occluded-region hallucination and visible-region alignment, SSGLD routes semantic conditions to Cross-Attention and spatial conditions to Self-Attention. As the spatial pathway’s core, we design the Spatial Aggregation Attention (SAA) module, which repurposes Self-Attention layers for pose-driven cross-image texture alignment using only sparse 2D skeletons, thereby eliminating the dependence on dense priors. On the standard DeepFashion benchmark, SSGLD outperforms comparable sparse-skeleton methods across all paired metrics and achieves texture fidelity competitive with dense-prior approaches. Ablation studies confirm the complementarity of the two routing pathways and the primary contribution of SAA, while qualitative results demonstrate improved robustness to severe occlusions, large pose variations, and dense-prior estimation failures. Together, these results show that SSGLD effectively resolves the trade-off between texture fidelity and pose-variation robustness in high-resolution paired PGPIS under sparse 2D skeleton conditioning.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.