Skip to content

Constructing embeddings for textual causality via representation regularization and semantic metrics

Aug 2026 · Knowledge and Information Systems · Vol 68 · 0 citations · 69 references

TL;DR

Compared to end-to-end classification models, this method of constructing an embedding space based on cross-level semantic metric learning enhances the ability to learn textual causality features and significantly reduces the dependency on annotated data.

View source

Similar papers

Open access Aug 2026

From Word Embeddings to Semantic Projections: Interpretability and Context in Web-Scale Semantic Analysis

This paper revisits semantic projections and related count-based representations as interpretable directional semantic structures for semantic analysis in document corpora and web-based information environments and demonstrates that semantic projections effectively capture persistent contextual structures while remaining sensitive to corpus-specific discourse communities.

Mabel López-Bordao, Antonia Ferrer-Sapena, Pablo Lara-Navarra et al. · 0 citations
Preprint Aug 2026

Logical Embeddings for Argument Analysis

It is proved that logical embeddings encapsulate the logical semantics of an argument, allowing for a better representation of its meaning, and that this encoding is optimal, in the sense that no logical information is lost in the process.

Leander Heldring, S. Torres · 0 citations
#artificial intelligence Preprint Aug 2026

Do General NLP Embeddings Capture Ontological Reasoning?

AVA is introduced, a systematic framework for evaluating whether embeddings distinguish logic-sensitive relational semantics in ontologies and knowledge graphs, and reveals a persistent gap between linguistic representation learning and ontology-level discrimination, challenging the assumption that strong NLP benchmark performance translates to Semantic Web competence.

Hamed Babaei Giglou, Jennifer D’Souza, S. Auer · 0 citations
Book Open access Aug 2026

Adaptive Pseudo-Labeling via Word Coherence for Topic Modeling

Topic modeling discovers latent semantic structures from document collections, providing interpretable insights applicable across a wide range of domains. However, conventional topic modeling approaches are limited by the absence of labeled data, requiring unsupervised learning for training. To overcome this challenge, we present adaptive pseudo-labeling for topic modeling (APT), a self-supervised learning framework designed to alleviate the need for labeled data. Our framework employs document embeddings derived from pretrained transformers and reconstructs the Bag-of-Words (BoW) representation by directly learning the semantic relationships between documents and words. Simultaneously, APT dynamically generates adaptive pseudo-labels to enhance topic coherence, leveraging word coherence extracted from the BoW representation and semantic relationships among document, word, and topic embeddings. On this basis, we integrate proxy-based deep metric learning into topic modeling to improve semantic coherence and diversity across topics. Accordingly, APT derives latent topics and document representations based on the distances between embeddings in the semantic space. Comprehensive experiments on benchmark datasets demonstrate that our APT framework outperforms conventional topic modeling approaches.

Bohan Yoon, Hyejin Jang · 0 citations
Conference 2026

Ambiguity-Aware Keyword-Enhanced Label-Aware Semantic Fusion for Text Classification

The rapid growth of textual data has made text classification a fundamental task in Natural Language Processing (NLP). However, real-world texts often exhibit semantic ambiguity, limited contextual information, and unclear category boundaries, which hinder conventional models from learning discriminative representations. To address these challenges, this paper proposes an Ambiguity-Aware Semantic Fusion Framework (AAK-LASFNet) for robust text classification.The proposed model constructs dual-view semantic representations by combining local contextual features extracted by a TextRCNN encoder with global semantic knowledge obtained from a large language model. An ambiguity estimation module is introduced to model semantic uncertainty, improving the model’s ability to handle ambiguous samples. Meanwhile, a label-aware attention mechanism and a keyword enhancement module are employed to strengthen category-related semantic cues. To further capture complex interactions between local and global representations, a high-order semantic fusion strategy is developed. In addition, a semantic consistency loss is imposed to align different semantic views and enhance representation stability.Extensive experiments on four benchmark datasets demonstrate that the proposed framework consistently outperforms strong baselines in terms of Accuracy, highlighting its effectiveness in alleviating semantic ambiguity in text classification.

Peijun Xie · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.