Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
The findings indicate that the evaluated VLMs exhibit a consistent behavioral limitation when processing typographic structures containing multiple spatial frequency layers.
Mert İncidelen, Yamen Kashkash, A. Berker et al.
· 0 citations