Beyond CLIP: A Critical Analysis of Representational Misalignment When LLMs Replace Text Encoders in Diffusion-based Image Generation
Abstract
Recent text-to-image diffusion systems have begun replacing CLIP and T5 text encoders with decoder-only large language models (LLMs), motivated by their stronger language understanding. This substitution, however, does not straightforwardly improve image-text alignment: naively using an LLM as the prompt encoder can substantially degrade prompt-following ability, an effect traced to a mismatch between how LLMs are trained and represent text and what diffusion conditioning mechanisms require. This paper presents a critical analysis of this representational misalignment, decomposing its causes into three categories: a training-objective mismatch between next-token prediction and the discriminative representations diffusion conditioning expects; a representational-geometry mismatch arising from the causal, positionally biased structure of decoder-only architectures relative to bidirectional encoders; and a granularity mismatch between LLM-encoded reasoning-level semantics and the spatially groundable signal that cross-attention and joint multimodal attention require. We audit existing mitigation strategies, including adapter-based bridging and deep-fusion architectures, against this taxonomy, and conclude with design implications for which mismatch sources are likely correctable through adapters and which may require rethinking the conditioning mechanism itself.