This work presents the first system that, from a single illustration, generates all the structured information a Live2D runtime consumes: ordered RGBA layers, a deformation mesh per layer, and the parameter-driven keypose vertex offsets that make the character move.
Abstract
Live2D is the dominant 2D character-animation format for anime characters and virtual avatars, representing each character as a stack of RGBA layers driven by per-layer mesh deformation. Despite its wide use in virtual streaming, mobile games, and interactive characters, authoring a Live2D model still demands weeks of manual layer separation, occlusion completion, mesh placement, and keyframing, and no prior generative method produces such a structured asset end-to-end. We present the first system that, from a single illustration, generates all the structured information a Live2D runtime consumes: ordered RGBA layers, a deformation mesh per layer, and the parameter-driven keypose vertex offsets that make the character move. Stage 1 casts layered decomposition as a layered diffusion process under a Live2D-aware organ-level taxonomy, producing an ordered RGBA stack with hidden-region completion. Stage 2 builds a content-conforming triangle mesh for each layer from its alpha channel alone, then predicts the keypose displacement field of all layers jointly: every vertex of every layer is one token, self-attention spans layer boundaries, and each displacement is factorised into a bounded direction and a log-magnitude. Joint rather than independent prediction is what makes the result a coherent character instead of separately plausible parts, and is our largest gain; scaling the network 112x yields none. On 50 held-out characters, under true generation with no teacher forcing, Stage 2 attains a per-vertex direction cosine of 0.768 (median 0.828). Because a layer's mesh derives from its alpha channel, a clothing layer can be re-textured from a natural-language instruction while the mesh and predicted animation are reused byte-for-byte. We further contribute Live2D-Bench, the first standardized benchmark for the task, and an 8,884-model Live2D corpus with layer and animation supervision.
The high photorealism and rendering efficiency of 3D Gaussian Splatting make it a promising approach for realistic 3D head avatar synthesis, leading to substantial progress in static avatar generation in recent works. However, current generation methods lack support for expression-driven animation, exhibit limited geometric and textural detail, and offer poor editability, which limits the practical applicability of this technology. Therefore, we propose AEAvatar, an Animatable and Editable Generative Gaussian Head Avatar framework for producing high-quality 3D avatars. AEAvatar employs a 2D GAN to generate per-pixel attribute maps, which are indexed via UV coordinates to produce 3D Gaussian attributes. The UV space is partitioned into static and dynamic regions, each used to index a separate set of Gaussians. Dynamic Gaussians are bound to the FLAME mesh to support facial deformation and animation, while static Gaussians represent rigid regions such as hair and shoulders. Thanks to its efficient feature mapping design, AEAvatar achieves high-quality animation at a resolution of \(512^{2}\) using only around 120K Gaussian primitives. This compact representation, coupled with an optimized caching mechanism, significantly boosts rendering efficiency, enabling real-time animation at up to 415 FPS. In addition, identity editing such as 3D face swapping can be easily performed by replacing either the dynamic or static set of Gaussians. These capabilities position AEAvatar as a practical and effective solution for synthesizing high-fidelity, real-time, and editable 3D head avatars. Project page: https://baoachun.github.io/AEAvatar/.
Achun Bao, Jie Guo, Xiu Li· ACM Transactions on Multimed...· 0 citations
StateFlow is presented, a state-centric framework for generative previsualization that uses an editable 3D world to organize scene structure, evolution, and cameras, while off-the-shelf video models enhance visual quality when higher fidelity is desired.
Yuyang Yin, Zixiang Li, Longxuan Deng et al.· 1 citation
This work shows that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass, enabling interactive 4D-controllable streaming generation for the first time.
Shiqian Li, Chenguo Lin, Zhi-Guang Liu et al.· 0 citations
This work presents a training-free multi-agent system that edits existing 3D meshes directly in Blender by emulating the iterative workflow of human artists, and views this work as an exploratory step toward visual-centric agentic geometry editing in professional graphics software.
Bo Pang, Jiaqi Pan, Xiao-Chen Zhang et al.· 0 citations
Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spatio-temporal distributions and the scarcity of 4D training data. We present AnimateAnyMesh++, a feed-forward framework for text-driven animation of arbitrary 3D meshes with substantial upgrades in data, architecture, and generative capability. First, we expand the DyMesh-XL dataset by mining dynamic content from Objaverse-XL, increasing the number of unique identities from 60K to 300K and substantially broadening category and motion diversity. Second, we redesign DyMeshVAE-Flex with power-law topology-aware attention and vertex-normal-enhanced features, which significantly improves trajectory reconstruction, local geometry preservation, and mit igates trajectory-sticking artifacts. Third, we introduce archi tectural changes to both DyMeshVAE-Flex and the rectified flow (RF) generator to support variable-length sequence training and generation, enabling longer animations while preserving reconstruction fidelity. Extensive experiments demonstrate that AnimateAnyMesh++ generates semantically accurate and tem porally coherent mesh animations within seconds, surpassing prior approaches in quality and efficiency. The enlarged DyMesh XL, the upgraded DyMeshVAE-Flex, and variable-length RF to gether deliver consistent gains across benchmarks and in-the-wild meshes. We will release code, models, and the expanded DyMesh XL at https://github.com/JarrentWu1031/AnimateAnyMesh-pp upon acceptance of this manuscript to facilitate research in 4D content creation.
Zijie Wu, Chaohui Yu, Fan Wang et al.· IEEE Transactions on Pattern...· 2 citations· ⚡1
D, a reference-guided renderer that extends Wan2.2 camera control from Plucker rays alone to a joint camera-plus-geometry interface and projects a neural 4D G-buffer from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior, supporting tracking+world-position correspondence as a practical 4D rendering condition.
Junhao Chen, Mingjin Chen, Henghaofan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.