Beyond CLIP: A Critical Analysis of Representational Misalignment When LLMs Replace Text Encoders in Diffusion-based Image Generation
Recent text-to-image diffusion systems have begun replacing CLIP and T5 text encoders with decoder-only large language models (LLMs), motivated by their stronger language understanding. This substitution, however, does not straightforwardly improve image-text alignment: naively using an LLM as the prompt encoder can su...