Skip to content
Conference

Hierarchical multi-scale cross-attention learning for blind image quality assessment

Sep 2026 · International Conference on Internet of Things, Communication Engineering, and Artificial Intelligence · Vol 14373, pp. 143730S - 143730S-6 · 0 citations · 9 references
Engineering

TL;DR

A novel Hierarchical Multi-Scale Cross-Attention Network that effectively captures both local distortion patterns and global semantic information for quality prediction and exhibits superior generalization capability compared to existing approaches is proposed.

Abstract

Blind image quality assessment (BIQA) remains a challenging task in computer vision due to the absence of pristine reference images. This paper proposes a novel Hierarchical Multi-Scale Cross-Attention Network (HMCANet) that effectively captures both local distortion patterns and global semantic information for quality prediction. The proposed architecture integrates a Swin Transformer backbone with a novel Cross-Scale Attention Module (CSAM) to aggregate multi-level features hierarchically. Additionally, a Distortion-Aware Feature Enhancement (DAFE) block is introduced to amplify quality-relevant representations while suppressing irrelevant noise. A composite loss function combining mean squared error and rank-aware loss is designed to improve prediction accuracy and monotonicity. Extensive experiments are conducted on four benchmark datasets, including LIVE, CSIQ, TID2013, and KADID-10k. The experimental results demonstrate that HMCANet achieves state-of-the-art performance with SRCC values of 0.971, 0.954, 0.923, and 0.917 on the respective datasets. Furthermore, ablation studies are performed to validate the effectiveness of each proposed component. The cross-dataset evaluation results indicate that the proposed method exhibits superior generalization capability compared to existing approaches.

View source

Similar papers

Open access Sep 2026

Attention Mechanism Guided Content-Aware No-Reference Image Quality Assessment

No-reference image quality assessment (NR-IQA) quantifies image distortion. It plays an important role in computer vision. Distorted images vary greatly in content. Many existing methods tend to fuse content information with quality prediction. However, they often overlook human visual perception. To address this issue...

Guo-Hong Zhou, Long-Sheng Wei · 0 citations
Open access Sep 2026

A Swin transformer-based framework for digital media image quality assessment

No-reference image quality assessment (NR-IQA) is an important task in image processing and is essential for automatically monitoring image quality during content distribution. Images captured under uncontrolled conditions may contain multiple authentic distortions, and their perceived quality depends on multiple dimen...

Xiao-Meng Xia, Jia Yong, Yi-Biao Long et al. · 0 citations
Open access Sep 2026

EDAFNet: efficient dual attention fusion network via multi-exposure image for HDR reconstruction

High Dynamic Range (HDR) reconstruction from multi-exposure Low Dynamic Range (LDR) images requires recovering a wide luminance range while preserving details in bright and dark regions under motion and exposure misalignment. High reconstruction fidelity, however, often comes with increased computational complexity. Th...

Ian Oliveira Teixeira, Q. Leher, Josue Lopez-Cabrejos et al. · 0 citations
Open access Aug 2026

Deep Edge-Aware Post-Processing for JPEG Enhancement: CNN-Based Artifact Reduction and Image Quality Restoration

A CNN-based edge-aware artifact reduction framework (CNN-AR) is proposed that integrates an enhanced deep super-resolution (EDSR) backbone with a holistically nested edge detection (HED) guided loss, enabling superior artifact suppression while preserving fine structural details.

Nupur, Nishant Kumar, Sajal Suhane et al. · 0 citations
Conference Sep 2026

Full-reference image quality assessment based on LS convolution

This study proposes a full-reference image quality assessment (FR-IQA) algorithm, named LSCNN, which conforms to the “large perception, small aggregation” characteristic of the human visual system. The LSCNN algorithm extracts important features that conform to human subjective perception by replacing some convolutiona...

Huang Lao, Si-Si Fan, Yu-Le An et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.