Prompt Yourself: Awakening Textual Semantics in 1D Visual Tokenizers
VLTok is a novel 1D hybrid tokenizer that unifies V isual and L anguage representations in a shared Tok en space through a self-prompted training paradigm, and achieves state-of-the-art performance in both image reconstruction and image generation.
Hualiang Wang, Siming Fu, Wei-Nan Jia et al.
· 0 citations