With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study seman...
Yu Zhao, Jia-Rui Wang, Hui-Yu Duan et al.· 0 citations
Aligning with the human visual system~(HVS) in perceiving and evaluating the quality of visual signals is a central objective of machine-vision-based visual quality assessment systems. With the rapid progress of large multi-modal models~(LMMs), visual question answering provides a promising paradigm for building unifie...
Zi-Heng Jia, Zi-Cheng Zhang, Jia-Ying Qian et al.· 0 citations
AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring the importance of effective quality assessment for such videos. To this end, we extend our previous dataset HVEval with pairwise preference a...
Sijing Wu, Yunhao Li, Huiyu Duan et al.· IEEE transactions on circuit...· 4 citations
MIE-Bench is introduced, the first large-scale multiple image editing benchmark with fine-grained human preference annotations and MIEScore, a multimodal large language model (MLLM)-based evaluation model enhanced with skill optimization and multi-dimensional supervised fine-tuning, to provide human-aligned feedback fo...
Zi-Tong Xu, Huiyu Duan, Xinyu Zhang et al.· 0 citations
This model leverages different layers of the large language model (LLM) decoder, treating them as multiple detectors that perform synchronous distortion detection using multi-level features, which helps mitigate the ambiguous foreground-background separation commonly encountered in the S2A problem.
Zi-Heng Jia, Yingji Liang, Jia-Ying Qian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.