Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason to believe the true task geometry has this property. The task-sensitive geometry of the feature space is given by the pu...
Andrew Bond, Ege Erdem Ozlu, Tuna Çimen et al.· 0 citations
Whole-slide pathology images (WSIs) contain gigapixel-scale visual content, creating a major scalability challenge for slide-level multimodal large language models (MLLMs). Existing approaches process thousands of patch tokens and typically apply compression only after slide encoding, leaving multimodal attention compu...
Ali Kerem Bozkurt, Baris Cem Bakay, Ibrahim Kulac et al.· 0 citations
BeyondMasks reframes video object removal as causal scene consistency rather than local reconstruction and provides a unified framework for its evaluation, and proposes CORE, a structured vision language model based evaluation protocol that jointly measures object disappearance and after effect consistency, aligning mo...