Skip to content
Open access

Semantic core identification as a method to overcome textoidness

Unknown authors
Jun 2026 · RESEARCH RESULT Theoretical and Applied Linguistics · Vol 12 · 0 citations

TL;DR

There is no reliable method to assess and achieve global semantic coherence in AI-generated translations, so this study aims to lay the foundations of a linguistic method for overcoming textoid-quality of machine translation results by means of semantic core identification.

Abstract

We believe that neural machine translation results intended to function as a text always have enough potential for a semantic core (i.e. a communicative center with text-forming properties) to be found and verbalized. The relevanceof this article is provided by two factors. On the one hand, machine translation software is widespread, easily available, and in active use; on the other hand, machine translation results have to be post-edited to the quality of a communicative text due to systematic disruption of its intra-textual connections in the machine translation results which turns out to be, in its raw, non-edited version, a set of separate sentences, in other words – a ‘textoid’ that should be fixed by an editor to function as a coherent text. Although frequent cross-checking between the original text and its translation helps eliminate occasional semantic errors and inaccuracies, the AI output in general still looks like a poor-quality text with a ‘machine DNA.’ This brings us to the core problem: now, there is no reliable method to assess and achieve global semantic coherence in AI-generated translations. That is why our study aims to lay the foundations of a linguistic method for overcoming textoid-quality of machine translation results by means of semantic core identification. Through a comprehensive approach that comprises such methods as abstraction, analysis, classification, synthesis, modeling, and measurement this study has achieved the following results: (a) a unique tool for semantic core identification was proposed relying on such well-known linguistic concepts as subject, predicate, and object, as well as on a basic subject-logical typology of semantic relations; (b) a need to adjust the initial core wording/formula was demonstrated in 46 % of cases; (c) the median core volume (31 %) in a textoid was determined for medical news; (d) basic principles of linguistic annotation (how to label specific linguistic, structural, or semantic features) were proposed as well as a system of notations; (e) a principle for representing the semantic core by means of graphic formulae was proposed for illustrative purposes; (f) ways for further scientific research were outlined. Conclusion: 52 textoids were analyzed to demonstrate applicability of our method, intended to serve as a reliable linguistic tool for identifying a semantic core which, in its turn, can function as (1) a text-forming essence that can be used in converting a textoid into a text; (2) a subject-logical benchmark for controlling and verifying translation, both for specific segments of the machine translation and for the text as a whole; and (3) a tool for interpreting unclear or contradictory passages within the textoid (without direct need to check up with the source text).

Read PDF

Similar papers

Open access Jul 2026

Texts Generated by Artificial Intelligence: Structure and Semantics

It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology.

L. Kravets, Viktória Stefuca, N. Libak et al. · 0 citations
Open access Sep 2026

The Shortcomings of Natural Language Processing (NLP) Models and Their Applications in Achieving Accuracy in Translating from Arabic to English

Translation is defined as the process of transferring meaning from one language to another. It is an extremely difficult and complex process because it involves not only transferring words but also ideas, culture, linguistic customs, and meanings derived from syntactic elements, their arrangement, word structure, and derivation. This is especially true in languages with complex structures, such as Arabic, which is characterized by its multiple linguistic contexts. These contexts have not been adequately addressed by NLP (Natural Language Processing) applications in machine translation models due to the lack of diverse contexts where metaphor, figurative language, grammatical inflections, morphological patterns and their connotations, and sentence structure all play pivotal roles in determining meaning. Furthermore, spoken language, with its inherent phonetic and expressive characteristics, conveys the text into broader semantic spaces. These spaces are influenced by the effect of intonation on specific syllables, the speaker's psychological state, the listener's mood, and accompanying body language, which transforms meaning into other subtle details. All of this, and more, is absent from machine translation, no matter how hard its creators try to imbue it with human emotions and feelings. This study highlights the importance of integrating in-depth linguistic analysis, contextual semantic modelling, and cultural awareness into natural language processing-based translation systems. By combining traditional linguistic insights with computational methods, the research offers a framework that can contribute to improving the accuracy of machine translation from Arabic to English.

Hilal Abdul-Raziq Sadiq, Zaxid Maxmudovich Islamov, R. Matibaeva et al. · 0 citations
Open access 2022

Semantic analysis of texts in a non-linguistic university based on Bloom’s taxonomy

Methods of semantic analysis and the appeal to the semantic side of the language based on Bloom’s taxonomy in the Russian language classes at a non-linguistic university in the study and analysis of the text are actual. The role of Bloom’s taxonomy in the step-by-step implementation of the content analysis algorithm and its appeal to it lies in its consistency, i.e., in the mandatory stages of conscious learning, such as knowledge, understanding, application, analysis, and synthesis. The use of this method contributes not only to the processes of memorizing and reproducing facts, but also allows you to establish a connection with the previously received information, generalize, change, interpret and transform it.In this regard, the text is an invaluable source of factual material that provides options for the diverse use of language units at various levels. In Russian language classes, the most effective methods for semantic text analysis are the associative method and the content analysis method. Tasks based on associations arouse students ‘ cognitive interest, and content analysis contributes to better assimilation of texts of different speech styles, including texts on the specialtySemantic analysis of texts based on Bloom’s taxonomy in a non-linguistic university contributes to the development of critical and creative thinking of students by forming the ability to determine the value of an idea based on a critical and objective consideration of arguments

R.D. Darkembayeva, N. Ozekbayeva, F. Sametova · 0 citations
Open access Jul 2026

PERSONAL NAME-BASED NEOLOGISMS IN MACEDONIAN ONLINE JOURNALISTIC DISCOURSE

Languages are permanently expanding their vocabulary with new words, which is an inevitable and constant process. The new words can be produced by the means of one language, borrowed from a foreign language, or created as a combination of one’s language means and borrowed elements. These novel words appear for the first time mainly in the journalistic articles, and then, depending on many factors, either spread into the general lexicon, or remain a part of a specialized language. The research question relates to the personal name-based neologisms in the Macedonian journalistic discourse aiming to reveal the most frequent word formation devices journalists use when creating such neologisms in their text. Thus, 2000 online news articles, informative and opinion based, are selected by using personal name-based neologisms as key words on Google Search. They serve as a sample, which is purposeful and cohesive. The data analysis rests upon the general qualitative inductive method and is enriched by the coding, i.e. organizational boxes are established in which the data are sorted for further analysis. The research results showcase that Macedonian journalistic discourse contains personal name-based neologisms which support the view that they are a common phenomenon in many modern languages. Further, the findings reveal that the suffixation is the most frequent language means used by journalists when generating personal name-based neologisms which reinforces previously stated outcomes. Furthermore, from the results it is obvious that the new personal name-based words are created with prefixes and that there are new compounds as well. In addition, these answers unveil that the personal namebased neologisms assist in achieving stylistic expression of the journalists’ text.

Violeta Janusheva · 0 citations
Open access Jul 2026

Evaluating the limits of machine translation for poetry: a multidimensional framework

The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment.

Beatriz Ribeiro Borges, P. H. R. Gabriel, E. Faria · 0 citations
Conference Open access 2026

The Translation of Body Part Idioms Using Chatgpt: a Comparative Analysis with Official Dictionaries

: In the era of digital transformation and the rapid advancement of generative artificial intelligence, the translation of idiomatic expressions has become a crucial benchmark for evaluating the cognitive and linguistic capabilities of Large Language Models (LLMs). This paper presents a detailed analysis of research conducted on a corpus of ten English body part idioms taken from the Pioneer B2 textbook used at Singidunum University. The aim of the research was to compare translations generated by the ChatGPT model with solutions from official idiomatic dictionaries, utilising Pavol Kvetko's classification and Mona Baker’s equivalence strategies as the theoretical framework. The analysis encompasses idioms of varying degrees of transparency, ranging from completely opaque to semi-idioms. The study results indicate a 90% accuracy rate in conveying meaning, alongside an unexpectedly high 60% correspondence of keywords in both languages. The research confirms that ChatGPT successfully identifies functional equivalents in the Serbian language, often prioritising the naturalness of expressions over literal translation. This work contributes to the discussion on the role of AI tools as assistants in translation and education, emphasising that while AI shows exceptional dexterity in mapping conceptual fields, human oversight remains essential for the final validation of stylistic nuances. The findings have significant applications for international scientific research, particularly in the domain of applying information technology in foreign language teaching.

Jelena Janackovic, Jovana Bošković, Jelena Mladenović · 0 citations