Expressivity In Multimodal Contrastive Learning
Hadamard-CLIP is proposed, which adds a single learned weight vector on top of the existing encoders and restores universal approximation of the joint for any number of modalities while preserving CLIP's fast, precomputable-embedding retrieval.
Andrew M. Stuart, Florian Wolf
· 0 citations