Skip to content

Author

Lingyu Zhu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Incorporating DINO Priors into Flow Matching for Low-Light Image Enhancement

Flow matching enables efficient low-light image enhancement (LLIE) with very few sampling steps, yet standard architectures lack explicit scene understanding, causing structural degradation and artifacts in challenging regions. We propose DINO-guided Flow Matching, which leverages a frozen DINOv3 backbone to provide illumination and structure priors for the Pixel MeanFlows framework. Specifically, we extract dual-layer features—shallow illumination-sensitive features and deep degradation-invariant structure features—and bridge the low-light/normal-light domain gap through a lightweight DINO Feature Corrector (DFC). The corrected features are injected into the flow-matching UNet via Retinex-inspired FiLM modulation and cross-attention, providing spatially adaptive guidance. Furthermore, we identify a systematic brightness drift problem arising from the marginal distribution mismatch between source and target domains, and address it with an Optimal-Transport Look-Up Table (OT-LUT) that pre-aligns the intensity distribution at negligible cost. Experiments on LOL-v2-real, LOL-v2-synthetic, and MIT-5K demonstrate state-of-the-art results in both distortion metrics (PSNR, SSIM) and perceptual quality (LPIPS).

Xiang-Rui Zeng, Ling-Yu Zhu, Jing-Ming He et al. · 0 citations
Aug 2026

Optimizing Fidelity-Perception Tradeoff via Large Vision–Language Model Prior for Image Compression

Current neural image compression (NIC) methods primarily focus on signal fidelity optimization. While perceptually optimized codecs can generate decoded images that better align with human visual preferences at equivalent bitrates, they raise authenticity concerns due to potential deviations from the original content. Therefore, achieving controllable decoding is crucial in various applications. This study presents a novel plug-and-play framework that leverages large vision-language model (LVLM) priors to balance fidelity and perception for existing NICs. Our approach consists of two key components: a scalable Low-Rank Adaptation scheme to controllably enhance the semantics of initially decoded images, and a two-stage agent-assisted decoding strategy with vision-language priors utilization. Specifically, the first stage extracts textual semantic information from an LVLM using decoded images enhanced by flexible fidelity-perception decoding, while the second stage effectively integrates semantic priors from LVLMs, further mitigating decoding semantic uncertainty and achieving higher-quality decoding. Extensive experiments on multiple benchmark datasets demonstrate that our method enables off-the-shelf NICs to achieve flexible control between optimal perceptual quality and signal fidelity.

Yudong Mao, Peilin Chen, Hao Luo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.