Skip to content
Conference

No-reference image quality assessment via multihead cross-scale attention fusion and contrastive ranking loss

Sep 2026 · Asia Conference onAsia Conference on Advances in Image Processing · Vol 14375, pp. 143750H - 143750H-8 · 0 citations · 16 references
Engineering

TL;DR

Three targeted enhancements to No-Reference Image Quality Assessment show state-of-the-art results on all three synthetic benchmarks and reveal configuration-dependent patterns on authentic-distortion data.

Abstract

No-Reference Image Quality Assessment (NR-IQA) must predict perceptual quality scores directly from distorted images without access to any pristine reference, making it substantially more challenging than full-reference approaches. While Transformer-based frameworks have recently advanced the state of the art, existing methods still leave room for improvement in cross-layer multi-scale feature interaction, the density of ranking supervision extracted from each training batch, and the spatial adaptability of the attention mechanism. Taking a code-reproduced TReS as the baseline, this paper introduces three targeted enhancements. First, a Multi-Head Cross-Scale Attention Fusion (MCSAF) module applies multihead self-attention to the concatenated feature sequences from different ResNet-50 layers, allowing each scale’s contribution to adapt to the input content instead of being determined by naive concatenation. Second, an Advanced Contrastive Ranking Loss extends TReS’s triplet loss—which exploits only the extreme-quality sample pair within each batch—into a joint formulation comprising all-pair dynamic-margin ranking, hard negative mining, and temperature-scaled InfoNCE contrastive learning; this raises the effective supervision density from 𝑂(1) to 𝑂(𝐵 2 ) with no additional inference cost. Third, a Deformable Attention mechanism with Learnable Position Encoding replaces fixed sinusoidal codes with end-to-end optimizable row-column embeddings and uses a lightweight offset MLP to predict 𝐾 sampling points per query, reducing attention complexity from 𝑂(𝑁2 ) to 𝑂(𝑁𝑀𝐾). Experiments on LIVE, CSIQ, TID2013, LIVEC, and KonIQ-10K show state-of-the-art results on all three synthetic benchmarks (CSIQ SROCC/PLCC: 0.9465/0.9593, +4.0%/+2.8% over baseline) and reveal configuration-dependent patterns on authentic-distortion data. Ablation studies consistently identify the Advanced Contrastive Ranking Loss as the most universally effective component.

View source

Similar papers

Conference Sep 2026

Hierarchical multi-scale cross-attention learning for blind image quality assessment

A novel Hierarchical Multi-Scale Cross-Attention Network that effectively captures both local distortion patterns and global semantic information for quality prediction and exhibits superior generalization capability compared to existing approaches is proposed.

Jiakuo Yan, Jun Zeng · 0 citations
Open access Sep 2026

Attention Mechanism Guided Content-Aware No-Reference Image Quality Assessment

No-reference image quality assessment (NR-IQA) quantifies image distortion. It plays an important role in computer vision. Distorted images vary greatly in content. Many existing methods tend to fuse content information with quality prediction. However, they often overlook human visual perception. To address this issue...

Guo-Hong Zhou, Long-Sheng Wei · 0 citations
Open access Sep 2026

A Swin transformer-based framework for digital media image quality assessment

No-reference image quality assessment (NR-IQA) is an important task in image processing and is essential for automatically monitoring image quality during content distribution. Images captured under uncontrolled conditions may contain multiple authentic distortions, and their perceived quality depends on multiple dimen...

Xiao-Meng Xia, Jia Yong, Yi-Biao Long et al. · 0 citations
2026

No-Reference Image Quality Assessment via Perception-Guided Distortion Representation Refinement

No-Reference Image Quality Assessment (NR-IQA) aims to predict perceptual image quality from distorted images without reference signals. Existing NR-IQA methods often incorporate visual attention through external weighting or feature fusion, while semantic and distortion cues are often not sufficiently organized for qu...

Yu-Tong Zhang, Si-Qi Zhou, Feng Liang et al. · 0 citations
Conference Sep 2026

Full-reference image quality assessment based on LS convolution

This study proposes a full-reference image quality assessment (FR-IQA) algorithm, named LSCNN, which conforms to the “large perception, small aggregation” characteristic of the human visual system. The LSCNN algorithm extracts important features that conform to human subjective perception by replacing some convolutiona...

Huang Lao, Si-Si Fan, Yu-Le An et al. · 0 citations
Open access Sep 2026

EDAFNet: efficient dual attention fusion network via multi-exposure image for HDR reconstruction

High Dynamic Range (HDR) reconstruction from multi-exposure Low Dynamic Range (LDR) images requires recovering a wide luminance range while preserving details in bright and dark regions under motion and exposure misalignment. High reconstruction fidelity, however, often comes with increased computational complexity. Th...

Ian Oliveira Teixeira, Q. Leher, Josue Lopez-Cabrejos et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.