Preprint
Aug 2026
Visual Distortion Detection in UGC Images Using Large Multimodal Models
This model leverages different layers of the large language model (LLM) decoder, treating them as multiple detectors that perform synchronous distortion detection using multi-level features, which helps mitigate the ambiguous foreground-background separation commonly encountered in the S2A problem.
Ziheng Jia, Yingji Liang, Jiaying Qian et al.
· 0 citations