Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Image tokenizers define the ``visual language''of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed...