Conference
Jul 2026
AttnFusion: gated cross-attention for fine-grained visual injection in latent diffusion models
This work presents a streamlined Attentive Fusion Network that employs text embeddings as queries to actively harvest pertinent cues from OpenCLIP-encoded visual data, and implements a zero-initialized, learnable Context Gating module that autonomously regulates the blending of both modalities.
Wu Fan
· International Conference on... · 0 citations