Aug 2026· Journal of ICT Research and Applications· 0 citations· 15 references
TL;DR
This approach addresses the data scarcity problem in Indonesian SRL by leveraging the availability of annotated English-language corpora by leveraging the availability of annotated English-language corpora.
Abstract
Semantic role labeling is a semantic analysis task that aims to identify the semantic relationships within a sentence, such as who did what to whom, where, when, and so on. Current semantic role labeling (SRL) models for the Indonesian language still face challenges in achieving strong performance due to the limited availability of annotated corpora, especially compared with English SRL models. Therefore, this paper develops an Indonesian SRL model using cross-lingual transfer. This approach addresses the data scarcity problem in Indonesian SRL by leveraging the availability of annotated English-language corpora. The method uses multilingual models and SRL datasets from both English and Indonesian. The multilingual models used in this study are XLM-R and mT5, both in base and large configurations. The datasets include Universal PropBank Indonesia and Gojali’s dataset for Indonesian, and CoNLL-2012 for English. Evaluation was conducted using test data from Universal PropBank Indonesia and Gojali’s dataset. Among all developed models, XLM-R large with cross-lingual transfer achieved the best performance, with an F1 score of 0.916 on Gojali’s dataset and 0.858 on the combined Universal PropBank Indonesia and Gojali datasets.
Cross-lingual information retrieval (CLIR) for low-resource regional languages remains challenging due to limited annotated training data, morphological complexity, and linguistic diversity. This paper systematically investigates Indonesian–Javanese CLIR by integrating traditional sparse lexical retrieval, multilingual dense retrieval, and advanced large language model (LLM)-based reranking within a multi-stage retrieval framework. To evaluate our approach, comprehensive experiments are conducted on the Indonesian–Javanese subset of the CLIRMatrix benchmark, comprising 20,000 Javanese documents, 1,000 Indonesian queries, and 11,000 graded relevance judgments. We compare BM25 with query and document translation, multilingual bi-encoder retrieval, and various LLM-based pointwise reranking methods. The experimental results show that BM25 combined with document translation provides the strongest initial baseline, achieving a Mean Average Precision (MAP) of 0.45, Mean Reciprocal Rank (MRR) of 0.71, and nDCG@10 of 0.55. While LLM-based reranking consistently improves weaker dense retrieval baselines, it often degrades search performance when applied directly on top of strong BM25-based retrieval. To address this, Reciprocal Rank Fusion (RRF) is introduced, which mitigates this degradation issue and achieves the best overall performance, reaching a MAP of 0.47, MRR of 0.74, and nDCG@10 of 0.57. These empirical findings highlight the critical importance of robust initial retrieval baselines and demonstrate that LLM-based reranking is most effective when combined with rank fusion strategies in low-resource CLIR settings.
Raden Mohamad Adrian Ramadhan Hendar Wibawa, Ika Alfina, Evi Yulianti· International Conference on...· 0 citations
This work introduces PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words, and conducts the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms.
Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al.· 0 citations
Low-resource languages remain challenging for cross-lingual semantic alignment because of limited parallel corpora. In addition, conventional symmetric alignment may distort the semantic space of a high-resource language through noisy low-resource updates. To address this issue, we propose Monolingual Anchoring for Cross-Lingual Semantic Alignment (MACA), focusing on Uyghur as a low-resource case study. MACA follows an asymmetric paradigm that treats the high-resource language as a fixed semantic anchor and transfers its semantic structure to the Uyghur side. The method consists of three components: (1) Anchored Embedding Initialization for newly introduced Uyghur subwords, (2) Cross-Lingual Neighborhood Anchoring for structural alignment between Uyghur and the anchor language, and (3) Monolingual Structure Anchoring for improving the internal semantic organization of Uyghur representations. Experiments centered on Uyghur-Chinese show that MACA outperforms LaBSE, the strongest off-the-shelf multilingual baseline in our comparison, by 7.67 points on cross-lingual STS. In an exploratory Uyghur-English zero-shot setting, MACA also surpasses LaBSE by 2.45 points without using Uyghur-English training data. These results provide evidence for the effectiveness of MACA in the evaluated Uyghur setting and suggest that monolingual anchoring may be further explored for related low-resource languages, such as Kazakh, Kyrgyz, and Uzbek.
Ruohao Yan, Huaping Zhang, Yuwen Niu et al.· Journal of King Saud Univers...· 0 citations
The section concludes with formal problem specification: given vocabulary V and corpus C, semantic analysis is formalized as a mapping problem preserving distributional properties, an optimization problem minimizing loss through gradient-based methods, and an evaluation problem assessing quality through semantic similarity, analogy, and downstream NLP task performance.
D. Akhmedjanova· Международный Журнал Теорети...· 0 citations
Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.
This article proposes a three-stage algorithm for training a model to recognize fake news in the Ukrainian language, using the multilingual transformer model XLM-RoBERTa, which solves this problem by utilizing cross-lingual knowledge transfer from English to Ukrainian.
Volodymyr Smahliuk, Ya. Kovivchak, Yu. Kynash· Big Data and Cognitive Compu...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.