Skip to content
Conference

Beyond CLIP: A Critical Analysis of Representational Misalignment When LLMs Replace Text Encoders in Diffusion-based Image Generation

Aug 2026 · 2026 International Conference on Secure Information Systems and Technologies (ICSIST) · pp. 1481-1489 · 0 citations · 25 references

Abstract

Recent text-to-image diffusion systems have begun replacing CLIP and T5 text encoders with decoder-only large language models (LLMs), motivated by their stronger language understanding. This substitution, however, does not straightforwardly improve image-text alignment: naively using an LLM as the prompt encoder can substantially degrade prompt-following ability, an effect traced to a mismatch between how LLMs are trained and represent text and what diffusion conditioning mechanisms require. This paper presents a critical analysis of this representational misalignment, decomposing its causes into three categories: a training-objective mismatch between next-token prediction and the discriminative representations diffusion conditioning expects; a representational-geometry mismatch arising from the causal, positionally biased structure of decoder-only architectures relative to bidirectional encoders; and a granularity mismatch between LLM-encoded reasoning-level semantics and the spatially groundable signal that cross-attention and joint multimodal attention require. We audit existing mitigation strategies, including adapter-based bridging and deep-fusion architectures, against this taxonomy, and conclude with design implications for which mismatch sources are likely correctable through adapters and which may require rethinking the conditioning mechanism itself.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.