Skip to content
Open access

Monolingual anchoring for low-resource cross-lingual semantic alignment: A case study on uyghur

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 32 references

Abstract

Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cross-Lingual Semantic Alignment (MACA), focusing on Uyghur as a low-resource case study. MACA follows an asymmetric paradigm that treats the high-resource language as a fixed semantic anchor and transfers its semantic structure to the Uyghur side. The method consists of three components: (1) Anchored Embedding Initialization for newly introduced Uyghur subwords, (2) Cross-Lingual Neighborhood Anchoring for structural alignment between Uyghur and the anchor language, and (3) Monolingual Structure Anchoring for improving the internal semantic organization of Uyghur representations. Experiments centered on Uyghur-Chinese show that MACA outperforms LaBSE, the strongest off-the-shelf multilingual baseline in our comparison, by 7.67 points on cross-lingual STS. In an exploratory Uyghur-English zero-shot setting, MACA also surpasses LaBSE by 2.45 points without using Uyghur-English training data. These results provide evidence for the effectiveness of MACA in the evaluated Uyghur setting and suggest that monolingual anchoring may be further explored for related low-resource languages, such as Kazakh, Kyrgyz, and Uzbek.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.