Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment
This work evaluates the capability of open-source Vision-Language Models to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset, and suggests that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation.