Spatial Super-Resolution and Temporal Interpolation of CCTV Video with ESRGAN and RIFE
Abstract
Conventional video surveillance cameras often capture sequences with low spatial resolution and reduced frame rates, which hinders scene analysis and forensic tasks. This paper evaluates a deep learning pipeline for the spatial and temporal enhancement of surveillance video that combines two existing pretrained models: ESRGAN for spatial super-resolution and RIFE for temporal frame interpolation. The contribution is the integration of both stages and its evaluation on CCTV footage under a reproducible degradation protocol, rather than a new architecture. Footage recorded with a commercial fixed surveillance camera at 848 $\times$ 480 pixels and 15 fps was degraded to $212 \times 120$ pixels following the standard Real-ESRGAN synthetic protocol. The pipeline is evaluated using PSNR, SSIM, and LPIPS, a learned perceptual image quality metric, benchmarked against bicubic upscaling. ESRGAN reaches an LPIPS of 0.311 against 0.714 for bicubic upscaling, while its PSNR is lower, 18.84 dB against 20.42 dB, a behavior consistent with the perceptiondistortion tradeoff. RIFE doubles the output frame rate from 15 to 30 fps, and its synthesized frames reach 37.12 dB in a hold-out test against 33.95 dB for frame averaging. The complete pipeline runs at 52.5 ms per input frame and 360 MB of GPU memory on a desktop card, which keeps it applicable to existing surveillance installations without replacing hardware.