It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figurat...