Jul 2026· Neural Networks· Vol 205 Pt A, pp.
109406
· 0 citations· 58 references
Medicine
TL;DR
A diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity and compares with recent image-to-image translation and diffusion-based baselines to observe competitive performance in both visual coherence and identity preservation.
Abstract
Facial sketch synthesis is important for cross-modal face analysis and digital forensics, yet existing models often suffer from structural distortions and identity inconsistency under limited paired training data. Traditional methods, primarily based on generative adversarial networks, often suffer from training instability and insufficient detail reconstruction. To address these limitations, this study proposes a diffusion-based framework with a stage-wise multi-condition guidance mechanism that enhances both structural and textural fidelity. Specifically, (i) during the downsampling phase, semantic segmentation features are fused to guide the model with accurate structural information, such as facial part locations; and (ii) during upsampling, a hybrid cross-attention mechanism is employed to integrate coarse image textures with denoised noise, refining fine-grained details. Additionally, we incorporate the Vision Transformer within the U-Net backbone to better capture global contextual information in low-frequency regions, further enhancing image realism. Experiments on multiple benchmark datasets demonstrate strong overall performance in SSIM and FSIM, indicating improved structural similarity and feature-level fidelity under the evaluated settings. Moreover, our method obtains competitive LPIPS results, indicating improved perceptual similarity. We further compare with recent image-to-image translation and diffusion-based baselines and observe competitive performance in both visual coherence and identity preservation.
Face photo-sketch translation is a significant task in cross-domain image generation. Traditional methods often struggle to balance global structure and local details, and they lack the ability of adaptive cross-domain feature fusion. To address these challenges, this paper presents a novel image generation method based on generative adversarial networks (GANs). In the early stage of the encoder, a Global-Local Fast Fourier Convolution module is introduced. The global branch employs Fast Fourier Convolution to capture long-range dependencies, while the local branch utilizes depthwise separable and standard convolutions to extract local textures. This parallel approach enables the simultaneous representation of global and local features. Additionally, a bi-directional gated channel attention module is incorporated to aggregate complementary information from both branches, thereby balancing overall consistency and fine detail realism. Furthermore, an Adaptive Cross-Domain Attention module is designed to dynamically select reference image features based on regional correlation: real details are preserved in highly correlated regions, while style alignment and fusion are performed in weakly correlated regions, effectively reducing interference from irrelevant features. Experimental results on the CUFS and CUFSF datasets demonstrate that the proposed method achieves superior visual quality outperforms or ranks second in terms of the LPIPS, FID, and FSIM metrics, validating its effectiveness in high-quality cross-domain image generation.
Lei Zhang, Houpan Zhou· International Conference on...· 0 citations
Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity degradation. We propose MTVDiff, a novel multimodal latent diffusion framework that synergistically integrates depth and textual information to address these limitations while preserving identity characteristics. The MTVDiff framework presents three core technical contributions: (1) a Dual-Branch Cross-Attention Fusion (DBCAF) module for multi-scale thermal-depth feature extraction and fusion; (2) a Gated Text-to-Visual Feature Alignment mechanism for semantically-guided generation; and (3) Spatial Feature Transformations (SFT) for adaptive multimodal prior integration. Extensive experiments on the MCXFace and SpeakingFaces datasets demonstrate that our multimodal approach significantly outperforms existing GAN-based and diffusion-based approaches across multiple metrics, achieving substantial improvements in both image quality and face verification performance, with FID reductions of up to 48.3% and Rank-1 accuracy improvements of up to 8.9\%. Our work provides a robust solution for face recognition systems operating under varying illumination conditions and advances the state-of-the-art in cross-spectral facial image translation through effective multimodal integration.
Zhiyuan Xia, Haojie Li, Jingyu Lin et al.· 0 citations
Image-based virtual try-on (VTON) has advanced rapidly with the emergence of high-resolution generative adversarial networks and diffusion-based synthesis models. However, many remaining failures, including garment misalignment, boundary artifacts, unrealistic deformation, and identity or body-shape distortion, are not caused only by generator limitations but also by the quality, structure, and fusion of upstream input representations. This paper presents an input-centric survey of image-based VTON systems. Unlike prior reviews that mainly organize the field by generative architecture, this work analyzes how geometry, semantic region control, garment conditioning, and multi-modal fusion shape the final try-on output. We review representative VTON methods, datasets, and evaluation practices, and group them according to the role of pose, body representation, parsing masks, garment appearance, and conditioning signals. We further discuss how representation errors propagate through warping, synthesis, and diffusion-conditioning stages. The survey highlights three main findings: (i) representation quality places an upper bound on synthesis realism, (ii) mask and semantic-region quality remain major bottlenecks even in recent diffusion-based approaches, and (iii) garment material representation is still weakly modeled in existing pipelines. Finally, we identify open research directions toward uncertainty-aware masks, material-informed garment embeddings, standardized evaluation protocols, and robust input fusion for real-world VTON deployment.
Le Thien Nhat Quang, Nguyen Van Hieu, Bui Cao Vu· 2026 11th International Conf...· 0 citations
Sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation, and a sketch encoder enhanced by Geometric Contour and Texture Detail preprocessing and MixStyle domain adaptation are proposed.
Baodong Wang, Fang Liu, Wei Cao et al.· The Visual Computer· 0 citations
Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascaded diffusion framework that introduces an explicit identity-conditioned semantic prior for multi-reference face inpainting. Our approach distills representative identity features from multiple references to reconstruct missing semantic regions, which then guide the diffusion process through a multi-stream conditioning architecture. This design provides strong semantic constraints when pixels are absent and stabilizes identity reconstruction while remaining compatible with prompt-driven edits. Experiments on CelebAHQ-IDI-5 and VGGFace2 demonstrate that ReSem-Face yields more reliable identity-preserving completion under severe semantic masks and improves text-controlled editing quality compared with representative baselines.
Feng Ding, Shuhuai Xie, Yue Zhou et al.· 0 citations