Preprint
Aug 2026
MMArt: A Multi-Perspective Multimodal Dataset for Visual Art Understanding
MMArt is introduced, a large-scale dataset of 74,234 WikiArt paintings, each annotated with four independently annotated perspectives plus a harmonized unified caption, produced by specialized vision-language models or human annotation and validated through complementary quality evaluations.
Shuai Wang, Wang-Yuan Ding, Yixian Shen et al.
· 0 citations