Skip to content

Semantic manifold-aware cross-modal similarity learning.

Aug 2026 · Neural Networks · Vol 205 Pt C, pp. 109541 · 0 citations · 53 references
Medicine

TL;DR

A Semantic Manifold-Aware Similarity Learning (SMSL) framework, where the term "semantic manifold" is used in an operational sense to denote a topology-aware organization of textual semantic units induced from token/entity interactions, rather than a strict low-dimensional differentiable manifold.

Abstract

Image-text matching (ITM) methods typically rely on positive-negative sample discrimination to learn cross-modal similarity. However, when similarity computation is primarily built upon sequence-level and linear representations, samples that are semantically similar but differ in underlying structural relations often remain ambiguously separated, resulting in blurred decision boundaries. To address this limitation, we propose a Semantic Manifold-Aware Similarity Learning (SMSL) framework, where the term "semantic manifold" is used in an operational sense to denote a topology-aware organization of textual semantic units induced from token/entity interactions, rather than a strict low-dimensional differentiable manifold. Specifically, the framework constructs a discrete semantic topology by disentangling intrinsic object, attribute, and relation dependencies within text, and injects the induced structural constraints back into the representation space through a topology injection mechanism, endowing textual embeddings with explicit topology awareness while preserving relational semantic continuity. The topology-aware textual representations are further exploited as semantic guidance to attend to and filter visual region features, reinforcing semantically relevant regions while suppressing redundant or distracting visual cues. After structure-aware enhancement on both the textual and visual sides, we introduce a dynamic threshold-based positive-negative decision boundary mechanism at the similarity computation stage. Unlike conventional fixed-margin strategies, this mechanism adaptively adjusts the decision boundary according to local semantic topology and cross-modal alignment certainty. In this way, the proposed method preserves the fundamental discriminative paradigm of ITM while shifting similarity evaluation from linear representation-level comparison to topology-aware similarity reasoning. Experiments on the Flickr30K and MS-COCO benchmarks demonstrate competitive performance in fine-grained ITM scenarios, validating the effectiveness of structurally informed and topology-aware similarity determination. The code is publicly available at: https://github.com/zhuliqi0309/SMSL.git.

View source

Similar papers

Preprint Aug 2026

Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information

A simple but principled approach, termed deep modality-shared self-expressive model (DeepMORSE), which discovers cross-modal structures via a modality-shared self-expressive model and simultaneously learns structured representations that conform to a union of modality-specific subspaces is proposed.

Xiang-Han Meng, Wei He, Zhi-Yuan Huang et al. · 0 citations
Open access 2026

GeoKR-Net: Geometry-Aware Representation Learning With Knowledge Anchoring for Sensitive Information Detection

With the increasing complexity of online communication, sensitive content frequently appears in variant, implicit, and context-dependent forms, posing significant challenges to conventional flat detection approaches. Domain-specific fine-tuning often induces representation anisotropy and semantic feature entanglement, leading to distorted semantic manifolds and unstable decision boundaries. Moreover, existing methods rarely consider the inherent ordinal dependencies among risk levels, limiting their reliability for fine-grained sensitivity assessment. To address these challenges, we propose GeoKR-Net, a hierarchically structured framework that jointly integrates geometric semantic refinement, adaptive knowledge anchoring, and hierarchical risk-aware inference. At the representation level, a geometric refinement operator performs isotropic correction and saliency-guided feature purification, reconstructing a balanced and discriminative semantic manifold. At the knowledge level, a dual-track adaptive anchoring mechanism explicitly models heterogeneous semantic evolution patterns through structured perturbation exploration and latent manifold expansion, enabling the transition from lexical matching to conceptual alignment while maintaining robustness under knowledge-cold-start scenarios. At the decision level, a hierarchical consistency constraint explicitly captures ordinal dependencies among risk levels, ensuring logically coherent and stable multi-level risk prediction. Extensive experiments conducted on four sensitive benchmarks and three general-domain datasets demonstrate that GeoKR-Net consistently outperforms strong baseline models. Further analyses of representation geometry and knowledge coverage verify its effectiveness in mitigating semantic anisotropy, enhancing risk-sensitive feature discrimination, and reducing reliance on static lexicons. The results highlight the importance of jointly modeling geometric structure, adaptive knowledge evolution, and hierarchical semantics for robust and interpretable sensitive information detection.

Kangyuan Qin, Min Yu, Ran Liu et al. · 0 citations
Open access 2026

Semantic-Aware Consistency Rectification for Cross-Modal Person Retrieval

Text-to-Image Person Retrieval (TIPR) faces significant challenges due to the “one-to-many” nature of cross-modal matching, where a single identity corresponds to multiple images with varying viewpoints. Standard metric learning approaches often enforce rigid alignment between a text description and all same-identity images, ignoring the semantic asymmetry between strictly paired and unpaired samples. This indiscriminate “hard-labeling” introduces optimization noise and fails to account for the lexical ambiguity inherent in natural language. To overcome these limitations, we propose the Dynamic Calibration and Robust Distillation Network (DCR-Net). Our approach introduces a Paired-Guided Consistency Rectification (PCR) module that utilizes strictly paired data as reliable anchors to dynamically calibrate the confidence of unpaired samples, effectively suppressing noise from inconsistent associations under mild-to-moderate corruptions. Furthermore, we design a Semantic Knowledge Distillation (SKD) strategy incorporating a Token-level Resilience Modeling (TRM) task. By leveraging soft probability distributions from a momentum teacher, this mechanism enhances the model’s robustness against linguistic variations and synonymous descriptions. Extensive experiments on benchmarks demonstrate that DCR-Net achieves state-of-the-art performance by exhibiting strong resilience against semantic noise and alignment rigidity in controlled settings.

Zi-Xuan Zhang, Lu-Ming Xiao · 0 citations
Conference 2026

Dual Associations Semantic Enhancement for Image-Text Matching

Image-text matching faces the significant challenge in effectively mitigating the large visual-semantic discrepancy among different modalities. Existing studies mainly address this by projecting multi-modal data into a common subspace for measuring their semantic similarities. However, most of them often overlook the inherent asymmetry for directional associations, i.e., the differences between im-age-to-text and text-to-image affinities, which often results in inaccurate retrieval results. To tackle the challenge, we propose a Dual Associations Semantic En-hancement (DASE) model to capture bidirectional image-text semantic associa-tions. Specifically, we first build a two-layer GCN fusion network to construct and mine the semantic associations for each modality. And then, due to the inher-ent asymmetry of directional associations, a Dual Associations Alignment Mod-ule (DAAM) is designed to capture the dual associations between visual and tex-tual modalities, enabling comprehensive cross-modal fine-grained interaction. Fi-nally, global alignment is incorporated with the local alignment to achieve full semantic matching across heterogeneous modalities in a unified embedding space. Experimental results on two publicly available datasets demonstrate that the pro-posed DASE model achieves significant performance improvements in image-text matching tasks compared to baseline methods, validating its effectiveness and superiority.

Xinlin Zhao · 0 citations
Jul 2026

Controlling Embedding Spaces with Text-Conditioned Transformations

Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification. These embeddings compress high-level semantics into a single vector, which comes at the cost of primarily expressing a dominant semantics like main object while suppressing other important attributes such as camera angle or color tone. We propose a text-conditioned transformation of visual embeddings that makes such attributes explicitly accessible. Given a natural language description of an attribute category (e.g.,"color"or"art style"), a network generates an affine transformation that emphasizes the specified attribute. Conditioning on text enables it to learn many attributes simultaneously, accessing them at inference time through an intuitive interface. The network is trained to align transformed embeddings with the frozen latent space, enabling retrieval using existing large-scale embeddings without any re-encoding. When applied to a full set, the same mechanism transforms the latent space for attribute disentanglement tasks such as multi-clustering. By operating directly in latent space, our method provides a unified and efficient framework for controlling embedding spaces, demonstrating state-of-the-art performance across both attribute-based retrieval and multi-attribute organization tasks with near-zero inference cost. Project page: https://joefioresi718.github.io/ControlEmbed_webpage/

Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani et al. · 0 citations
Jul 2026

DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

DINOde is proposed, an ODE-based framework that continuously aligns CLIP text embeddings with the DINO visual manifold through a continuous ODE trajectory, and achieves state-of-the-art performance across multiple OVSS benchmarks.

Sung-Hoon Yoon, Hoyong Kwon, Chang-Hwan Oh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.