Semantic Specificity Degradation in Zero-Shot Generative Vision-Language Models
Abstract
Vision-Language Models (VLMs) have shown strong zero-shot performance in generating free-form image descriptions. However, most evaluations focus on hallucinated content, while less attention is given to whether models preserve specific object identities. This study investigates whether zero-shot generative VLMs retain fine-grained semantic information or tend to produce broader category labels. We introduce the Generic Downgrade Rate (GDR), a metric that measures predictions that are correct at a general category level but incorrect at the specific class level. A controlled zero-shot experiment was conducted using BLIP-2 (OPT-2.7B) on a balanced subset of Food-101 containing 101 classes and 10 images per class (N = 1,010). Model outputs were evaluated using Fine Accuracy, Coarse Accuracy, and GDR. Fine Accuracy was 0.0099, while Coarse Accuracy reached 0.3069. The resulting GDR of 0.2970 indicates that many predictions captured the general food category but failed to preserve the specific class identity. Bootstrap confidence intervals showed that the difference between fine- and coarse-level accuracy was stable across resampled data. Overall, the results indicate a clear loss of semantic specificity in zero-shot generative outputs. By distinguishing general semantic correctness from exact class identification, this evaluation provides a more detailed view of how VLMs preserve different levels of semantic information.