Lightweight Probabilistic Downscaling from a Deterministic Base Model
This work finds that a two-stage training curriculum, combining deterministic pretraining with probabilistic tuning, transfers well to downscaling, beating the state-of-the-art for RMSE.