Sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation, and a sketch encoder enhanced by Geometric Contour and Texture Detail preprocessing and MixStyle domain adaptation are proposed.
Face sketch-to-photo synthesis plays a key role in computer vision with applications in law
enforcement, digital entertainment, and human–computer interaction. Existing generative
adversarial network-based methods typically face mode collapse, training instability, and
poor performance across different sketching styles...
Inas Ismael Imran, Ahmed A. Hashim· Advances in Artificial Intel...· 0 citations
This work proposes RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing, and introduces Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial region...
The proposed sketch-guided face generation pipeline based on a Conditional Variational Autoencoder designed to generate facial reconstructions from sketches and conditional attributes uses a stochastic preprocessing pipeline to extract edge maps from facial photographs, reducing dependence on manually paired sketch-pho...
Edson M. Odake, E. P. Ribeiro· Multimedia tools and applica...· 0 citations
This work shows that identity-conditioned face synthesis can be performed at a substantially lower computational cost by a latent Consistency Model with few iterations, without compromising image quality for large-scale synthetic face generation.
Tiago Kienen Chaves, Bernardo Biesseck, David Menotti· 0 citations
Automatic facial rigging across heterogeneous mesh topologies remains challenging because high-quality expression supervision is often tied to canonical templates, while deformation transfer to arbitrary meshes can introduce geometric artifacts and correspondence errors. We present TopoRig, a topology-agnostic facial r...
Andrew Fleet, Soroush Mehraban, Vida Adeli et al.· 0 citations
Although diffusion-based methods have substantially improved the controllability of multimodal face synthesis, their semantic alignment remains suboptimal because most existing approaches rely on implicit latent-space objectives to model the relationship between denoising variables and multimodal conditions. Such impli...
Yu-She Cao, Xue-Chao Zou, Xing Xi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.