This work proposes RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing, and introduces Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy.
Abstract
Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods primarily align global textual semantics with facial structures but often struggle to capture subtle local deformations, such as eyebrow tension, cheek contraction, and asymmetric mouth motions, resulting in limited geometric fidelity and editing precision. To facilitate fine-grained text-driven facial modeling, we first construct FaME-G2E, a large-scale multimodal dataset containing detailed text--mesh annotations and paired text--blendshape samples for unified 3D facial generation and editing. Based on this dataset, we propose RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing. Specifically, the Multi-Scale Retrieval Fusion (MSRF) module retrieves semantically consistent global and regional facial priors and fuses them in the blendshape space, suppressing conflicting local deformations while preserving coherent deformation patterns. Furthermore, we introduce Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy. Extensive experiments on FaME-G2E demonstrate that RAGMesh achieves superior performance over state-of-the-art methods in local geometric accuracy, text-guided controllability, regional editing precision, and inference efficiency. Video demo is available at https://youtu.be/Yr0_XkpWcNk, and the source code and dataset will be released upon paper acceptance.
Sketch-text-driven rectified flow (STDRF), a conditional rectified-flow framework for identity-preserving and semantically controllable 4D face generation, and a sketch encoder enhanced by Geometric Contour and Texture Detail preprocessing and MixStyle domain adaptation are proposed.
Baodong Wang, Fang Liu, Wei Cao et al.· The Visual Computer· 0 citations
Recent Gaussian Splatting–based head avatar reconstruction methods achieve photorealistic rendering by binding Gaussian primitives to 3D Morphable Models (3DMMs), enabling geometrically consistent deformation and animation control. However, the optimization process lacks explicit surface opacity constraints, causing in...
Junhee Cho, Seonghak Lee, Jongmin Lee et al.· IEEE Access· 0 citations
Face-Edit-Attributes, the largest collection of $169$ facial editing attributes focused on hair, accessories, and pose edits, is presented, which shows that most models performed hair and accessory edits well, but struggled with editing pose.
A deep learning framework is proposed to directly predict re-topologized facial meshes from synthetic multiview images generated with Visage Craft, an in-house physically based rendering system using an Appearance 3D Morphable Model (A3DMM).
The representation of 3D clothed humans as standardized 2D UV texture and displacement maps over an underlying body model has long been studied. This compact representation is enticing as it enables pretrained image networks to process, generate, and edit 3D avatars, but is only useful if scans are accurately aligned a...
Margaret Kostyrko, Yuxuan Xue, Garvita Tiwari et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.