A Text-Guided Cross-Modal Diffusion Framework With Attention-Based Hand-Drawn Sketches for Face Synthesis
Face sketch-to-photo synthesis plays a key role in computer vision with applications in law enforcement, digital entertainment, and human–computer interaction. Existing generative adversarial network-based methods typically face mode collapse, training instability, and poor performance across different sketching styles...