IREF-VITON: Input Representation Enhancement and Fusion in Image-Based Virtual Try-On Systems
Image-based virtual try-on (VTON) has advanced rapidly with the emergence of high-resolution generative adversarial networks and diffusion-based synthesis models. However, many remaining failures, including garment misalignment, boundary artifacts, unrealistic deformation, and identity or body-shape distortion, are not caused only by generator limitations but also by the quality, structure, and fusion of upstream input representations. This paper presents an input-centric survey of image-based VTON systems. Unlike prior reviews that mainly organize the field by generative architecture, this work analyzes how geometry, semantic region control, garment conditioning, and multi-modal fusion shape the final try-on output. We review representative VTON methods, datasets, and evaluation practices, and group them according to the role of pose, body representation, parsing masks, garment appearance, and conditioning signals. We further discuss how representation errors propagate through warping, synthesis, and diffusion-conditioning stages. The survey highlights three main findings: (i) representation quality places an upper bound on synthesis realism, (ii) mask and semantic-region quality remain major bottlenecks even in recent diffusion-based approaches, and (iii) garment material representation is still weakly modeled in existing pipelines. Finally, we identify open research directions toward uncertainty-aware masks, material-informed garment embeddings, standardized evaluation protocols, and robust input fusion for real-world VTON deployment.