Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems
Two state-of-the-art multimodal models, Gemma-3 and Qwen-VL, are assessed on their ability to interpret mechanical problem images by eliciting a step-by-step chain of thought (CoT) and a final answer, and final answers are compared to verified solutions to measure accuracy.
Henry Fordjour Ansah, Shreya Banerjee, Pranish Ghimire
· 0 citations