Skip to content

SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit

Aug 2026 · 0 citations · 36 references
Computer Science

Abstract

Lexical semantic change (LSC) is commonly modelled through vector-space representations, but these approaches often provide limited insight into which aspects of usage are changing. Diachronic corpus research instead examines interpretable dimensions such as syntactic behaviour, morphology, and constructional patterns, but typically through separate analytical workflows. We present SynFlow, an open-source toolkit for multidimensional diachronic analysis of linguistic usage. SynFlow converts linguistic observations into period-specific distributions and applies a shared workflow across dependency-based co-occurrences, morphological features, constructional configurations, and externally derived representations such as Frame Semantics. It supports different distance measures, together with value-level decomposition, statistical testing, and incremental clustering of lexical fillers. We demonstrate SynFlow through a qualitative case study of the German adjective viral, showing how a single semantic development is reflected across syntactic, lexical, constructional, and morphological dimensions. We further report previously published results on SemEval-2020 Task 1 to situate the performance of these representations relative to existing lexical semantic change detection systems.

View source

Similar papers

#natural language process... Preprint Sep 2026

Dynamics of meaning: Towards the Evaluation of Diachronic Semantic Change in Sinhala

Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and the limitations of static embedding alignments. This study investigates the diachronic evolution of the Sinhala language from the 13th to the 20th century using a multi-stage computational framework. We first align century-specific Word2Vec and FastText embeddings using Similarity Matrix Based Alignment (SMA) and Orthogonal Procrustes (OP) techniques, finding that OP alignment provides more stable neighbourhood tracking for identifying temporal similarity dips. To move beyond aggregate measures, we introduce a Bidirectional Semantic Impact Pruning approach using contextualised embeddings from a fine-tuned Llama-3.1-8B. By applying Leave-One-Out (LOO) diagnostics, we attempt to isolate influential sentences to distinguish between systemic semantic shifts and transient polysemic expansion. Our results show that semantic drift in the fine-tuned Llama-3.1-8B is not evenly distributed across all usages. Instead, a significant part of the change is driven by a smaller set of high-impact contextual instances, rather than gradual and uniform change across all occurrences. This work provides a preliminary framework for diachronic analysis in low-resource contexts, highlighting the trade-offs between model sensitivity and data availability.

Nevidu Jayatilleke, Nisansa de Silva · 0 citations
Open access Aug 2026

From Word Embeddings to Semantic Projections: Interpretability and Context in Web-Scale Semantic Analysis

This paper revisits semantic projections and related count-based representations as interpretable directional semantic structures for semantic analysis in document corpora and web-based information environments and demonstrates that semantic projections effectively capture persistent contextual structures while remaining sensitive to corpus-specific discourse communities.

Mabel López-Bordao, Antonia Ferrer-Sapena, Pablo Lara-Navarra et al. · 0 citations
Open access Aug 2026

Semantic analysis of problems in natural language processing and their mathematical interpretation

The section concludes with formal problem specification: given vocabulary V and corpus C, semantic analysis is formalized as a mapping problem preserving distributional properties, an optimization problem minimizing loss through gradient-based methods, and an evaluation problem assessing quality through semantic similarity, analogy, and downstream NLP task performance.

D. Akhmedjanova · 0 citations

Bridges Between Words and Senses

The proposed k-Multilingual Concept model allows to uncover novel layers of lexical knowledge in the form of multifaceted conceptual links between naturally disambiguated sets of words.

Francesca Grasso, Vladimiro Lovera, Luigi Di Caro · 0 citations
Open access Aug 2026

Research on the Mining and Visualization Analysis of Semantic Evolution Patterns of English Name Rotation Words Based on Corpus

Noun-to-verb conversion is a common form of English word-class conversion, and its semantic evolution reflects important mechanisms of lexical innovation. Existing studies often lack systematic mining of semantic evolution patterns and rely on static or single-dimensional visualization. This study constructs a corpus-based research framework that integrates corpus retrieval, semantic annotation, pattern mining, and interactive visualization. Typical noun-to-verb samples are selected from the British National Corpus, the Corpus of Contemporary American English, and the Oxford historical English corpus. Based on explicit semantic mapping rules, frame semantics, and dependency theory, the study identifies four core evolution patterns: metaphorical extension, metonymic mapping, semantic generalization, and semantic narrowing. Visualization using CiteSpace and ECharts presents pattern distribution, diachronic evolution trajectories, and semantic association networks. The results confirm the feasibility of combining corpus analysis with interactive visualization for semantic evolution research. The methodology can also support technical terminology tracking and semantic indexing in engineering corpora, including antenna systems, electromagnetic waves, and propagation-related texts.

Rui Zou · 0 citations
Open access Aug 2026

Construction of a Dialect-Sensitive Javanese Semantic Lexicon to Support Machine Translation Systems

The development of linguistic resources for natural language processing (NLP) in Javanese remains limited, especially regarding the representation of semantic relationships between different speech levels. This study aims to construct a Javanese semantic lexicon that integrates Indonesian lexical equivalents with three Javanese speech levels: ngoko, krama alus, and krama inggil. A research design based on lexical resource construction was employed, using a Javanese digital dictionary as the primary data source. The methodology included data extraction, preprocessing, semantic lexicon construction, analysis of speech level variation, and a preliminary exploration of polysemous lexical entries using automatic identification, followed by validation by native speakers. The resulting semantic lexicon successfully represents lexical relationships between levels in a structured manner. Analysis of speech-level variation revealed that partially distinct lexical patterns were the most dominant, with 733 entries, followed by fully distinct patterns (193 entries) and identical patterns (21 entries). These findings indicate that speech-level differences in Javanese are selectively realized and should be explicitly considered in the development of linguistic resources. Furthermore, preliminary exploration of polysemous candidates demonstrated that dictionary-based automatic identification can overestimate polysemy without linguistic validation. Only a limited number of lexical entries exhibited features consistent with genuine polysemous relationships. This study provides an initial basis for the development of Javanese semantic resources that are sensitive to speech-level variation and semantic complexity. The constructed semantic lexicon has the potential to support future research in NLP applications in Javanese, including politeness identification, lexical normalization, word sense disambiguation, and machine translation.

Musthofa Galih Pradana, Ridwan Raafi’udin, Nurul Afifah Arifuddin et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.