#artificial intelligence
Jun 2026
TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment
This work proposes TEVI, a framework that uses captions as a signal for what to retain from image embeddings and uses sparse autoencoders to disentangle image embeddings and train a masking module to selectively reconstruct the embedding based on a given caption.
S. Mahajan, Sukrut Rao, Jiahao Xie et al.
· arXiv.org · 1 citation