This paper presents the first application of visual autoregressive (VAR) next-scale prediction over a discrete image codebook to the RAW-to-sRGB ISP task and proposes a frequency-decomposed color loss that separately supervises low-frequency tone via wavelet LL cosine similarity and chromatic edges via detail-band $\ell_1$.
Abstract
RAW-to-sRGB image signal processing (ISP) must recover perceptually faithful colors and fine details from sensor measurements, often under imperfect spatial alignment and missing camera metadata. This paper presents, to the best of our knowledge, the first application of visual autoregressive (VAR) next-scale prediction over a discrete image codebook to the RAW-to-sRGB ISP task. We adapt a frozen 1.10\,B-parameter VAR backbone for RAW-conditioned ISP with only 32.93\,M trainable parameters (2.99\%), and propose a frequency-decomposed color loss that separately supervises low-frequency tone via wavelet LL cosine similarity and chromatic edges via detail-band $\ell_1$. On the Zurich RAW-to-sRGB benchmark, the method improves PSNR-Y from 21.31 to 21.89\,dB and reduces LPIPS from 0.276 to 0.218 on the full 1,204-image test set. Diagnostic experiments show that the VAR prior preserves structure well, but continuous color transfer remains the dominant bottleneck: oracle affine correction recovers 3.8\,dB, while learned color heads yield marginal gains.
SF-GAL, a prior-calibrated spatial-frequency Retinex decomposition framework for unsupervised low-light image enhancement, which calibrates a CLAHE-derived structural prior, decomposes low-light features through complementary spatial and wavelet branches, and uses structure-guided frequency modulation to regulate frequ...
Xin-Hua Dong, Yu Gao, Hongmu Han et al.· The Visual Computer· 0 citations
A novel tone mapping method that not only bridges the gap between HDR RAW inputs and the LDR sRGB requirements of detection networks but also achieves end-to-end optimization with downstream tasks and achieves real-time processing of 4K high-bit-depth HDR inputs on NVIDIA Jetson platforms.
Gongzhe Li, Linwei Qiu, Peibei Cao et al.· Neural Information Processin...· 2 citations
An enhancement pipeline that operates entirely within classical signal processing is proposed, providing a transparent alternative to black-box machine learning methods while remaining practical on standard personal computers.
Swarnajit Bhattacharya· Asian journal of applied sci...· 0 citations
I present a training-free low-light enhancement method that combines local bright-channel illumination estimation, Retinex division, and edge-preserving denoising. For a fixed illumination estimate, a conditional Negative -Binominal psueduo-count method characterises the heteroscedastic noise amplified by division. The...
This paper proposes BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation, and develops a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mi...
Tianyu Zhu, Ying Fu, He-Song Li et al.· IEEE Transactions on Pattern...· 0 citations
This work validates the design effectiveness of decoupling global and local representations within a frozen backbone, and establishes a new baseline for parameter-efficient enhancement.
Yan-Peng Cao, Yue Wang, Ming-Hui Liang et al.· Pattern Analysis and Applica...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.