From a Single Image to a 3D Model: A Review of 3D Human Reconstruction from Monocular Images
Abstract
In recent years, 3D human reconstruction technology has been widely applied across various industries such as digital humans, virtual reality, gaming, and film production. Among them, the task of generating a 3D human body model from a single two-dimensional image, known as single-image 3D human body reconstruction—poses significant technical challenges due to the relatively lack of input information. This article systematically reviews the mainstream technical routes and their advantages and drawbacks from 2015 to 2026. It starts from the early method of Skinned Multi-Person Linear Model (SMPL) parametric template, then progresses to the voxel prediction method represented by Deephuman, followed by the implicit function method centered on the Pixel-Aligned Implicit Function (PIFu) series as the benchmark, and finally concludes with the emerging neural rendering and real-time reconstruction technologies in recent years. The three significant experimental datasets, namely PIFu, Pixel-Aligned Implicit Function High-Resolution (PIFuHD) and Geometry and Pixel-Aligned Implicit Function (Ge0-PIFu), which have milestone importance, are analyzed using indicators such as chamfer distance, P2S and normal error. The improvement logic and applicable boundaries of each method are also discussed. Finally, the current practical challenges faced in this field, such as data deficiencies and insufficient social popularity, are summarized. The future technologies, like large model frontier inference and real-time rendering, are also prospectively discussed.