Enhanced Anime Image Generation by Comparative Analysis of DCGAN, LSGAN, and WGAN-GP
Abstract
: In recent years, anime-style image generation has become a prominent direction within generative adversarial network (GAN) research. However, a systematic exploration into the performance differences among various GAN architectures, specifically for anime face generation is still lacking. Therefore, this study utilizes the Anime Faces dataset and compares three classical GAN models-DCGAN, LSGAN, and WGAN-GP-under consistent data preprocessing and experimental conditions. Quantitative evaluation was performed using two widely recognized metrics, Fréchet Inception Distance (FID) and Inception Score (IS), complemented by a qualitative visual assessment of generated images. Results indicate that LSGAN achieves a balanced performance between training stability and image fidelity, producing images with superior detail and realism. DCGAN exhibits strong initial generation quality but struggles in later training phases due to optimization difficulties, resulting in fluctuating image quality. WGAN-GP, despite its theoretical advantages, performed inadequately on this style-consistent dataset, frequently encountering issues such as mode collapse. Consequently, this research recommends LSGAN as the preferred GAN architecture for anime face generation tasks. These findings provide valuable guidance for GAN model selection in stylized image generation applications.