Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI
It is argued that medical-image VLM evaluation should report verbalized-confidence reliability, confident error, hallucination, and abstention alongside accuracy, and that medical adaptation improves tumor-presence detection without improving confidence reliability.