Skip to content
Open access

Texts Generated by Artificial Intelligence: Structure and Semantics

Jul 2026 · Computer and Decision Making: An International Journal · Vol 3, pp. 932-946 · 0 citations

TL;DR

It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology.

Abstract

The article describes texts generated by artificial intelligence, with an emphasis on Ukrainian-language material, which remains understudied. The aim of the study was to identify specific identifying characteristics of AI texts by comparing their semantic, structural, and stylistic parameters with texts written by humans. The study combined systematic analysis and synthesis, a comparative approach, content analysis, and quantitative linguistic methods (Type–Token Ratio, syntactic complexity analysis by T-unit), as well as semantic modeling, fact verification, and semantic-stylistic analysis. It has been proven that Ukrainian-language AI texts formally meet the basic criteria of textuality (cohesion, coherence, articulation), which are implemented through the probabilistic combination of templates rather than the author’s cognitive and communicative activity. Typical markers of machine generation have been identified: template composition (introduction – main part (3–5 subtopics) – conclusion), homogeneous paragraphs of “average” length, predominance of direct word order, presence of passive constructions, excessive frequency of formal connectors, structural and lexical monotony, errors in word usage. Semantic analysis revealed a combination of formal correctness with factual “hallucinations,” low information density, a predominance of neutral style, superficial, statistically determined expressiveness, and emotional masking. It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology. The risks to linguistic norms, speech culture, and information security are emphasized, as well as the need to develop critical competence in users regarding the perception of AI content.

Read PDF

Similar papers

Open access 2026

Characterization and Mechanisms of Lexical Complexity in AI-Generated Texts: A Comparative Corpus-Based Study

: Based on a corpus-based methodology, this study analyzes the intrinsic reasons for the high level of lexical complexity observed in Artificial Intelligence Generated Content (AIGC). The research compares 24 English argumentative essays written by AI with 24 second-language (L2) learner essays reaching the IELTS Writing Task 2 Band 7 level. Under controlled conditions of identical genre and topic, the study performs quantitative statistics across three dimensions: lexical sophistication, semantic abstraction, and information density. Statistical results indicate that the frequency of advanced vocabulary in AI texts is significantly higher, approximately 2.3 times that of human texts. The proportion of abstract nouns reached 9.14%, far exceeding the 3.32% found in human texts, suggesting that AI expressions tend toward nominalization and conceptualization. Regarding overall information organization, the lexical density of AI texts was 69.9%, also surpassing the 60.3% of human texts, reflecting a stronger tendency for information condensation and phrasal structures. The analysis points out that the complexity of AI text primarily stems from its mechanism of selecting vocabulary based on probability distributions. This mechanism favors longer words, abstract nouns, and words with high semantic content, thereby forming a highly compact linguistic surface. Such complexity is essentially a formal feature at the statistical level and is not entirely equivalent to the proficiency levels corresponding to human L2 acquisition. These findings provide empirical references for AI text identification, the refinement of writing evaluation standards, and L2 writing pedagogy.

Lulu Chen · 0 citations
Open access Jul 2026

Syntactic Complexity in AI-Generated vs. Human-Authored Linguistic and Literary Texts

This paper examines how the syntactic complexity of academic writing is affected mainly by the source of authorship (AI-generated or human) or by the genre of disciplinary writing (linguistic or literary). The primary purpose is to test the syntactic-complexity differences between these variables and to establish the degree of influence of genre conventions on structural variation. The importance of the research is that it adds to the existing discussions about AI as a phenomenon in academic writing and, specifically, whether AI-generated texts are capable of syntactically reproducing the specific norms of a specific discipline of writing. To this end, the comparative corpus-based design was used. The sample consisted of 20 introduction sections: equal numbers of linguistic and literary texts and equal numbers of human-authored AI-generated texts. The Second Language Syntactic Complexity Analyzer (L2SCA) extracts fourteen syntactic complexity measures, which include length of production unit, subordination, coordination and phrasal sophistication. The results indicate that the complexity of syntax is genre-based and not source-based. Although no differences were found to be constant in both AI-generated and human academic introductions in linguistic data, there were much higher levels of subordination in literary texts of human origin. In general, the discipline genre had a more significant effect on syntax variation than the authorship source. The paper suggests the implementation of genre-sensitive models to assess AI-written academic texts and recommends additional studies that would use a bigger sample and discourse analysis.

Asia A. Alheety, Meethaq Khamees khalaf, Hussam J. Mohammed · 0 citations
Open access Aug 2026

LINGUISTIC FEATURES OF AI-GENERATED TEXTS AND METHODS FOR THEIR AUTOMATED IDENTIFICATION

The article examines linguistic features of Ukrainian socio-political news texts generated by large language models and methods for their automated identification. The aim is to substantiate lexical, stylistic, compositional and semantic markers that may indicate AI-generated text and to outline a computational-linguistic detection framework. The study proposes combining philological interpretation with NLP procedures: preprocessing, lexical diversity assessment, clustering, vectorization and transformer-based classification. It is argued that automated detection should not replace expert linguistic analysis but should serve as an auxiliary tool for evaluating the probable origin of a text in media linguistics, fact-checking, and educational practic

N. Babkova, D. Huliieva, Z. Kochuieva et al. · 0 citations
Open access 2022

Semantic analysis of texts in a non-linguistic university based on Bloom’s taxonomy

Methods of semantic analysis and the appeal to the semantic side of the language based on Bloom’s taxonomy in the Russian language classes at a non-linguistic university in the study and analysis of the text are actual. The role of Bloom’s taxonomy in the step-by-step implementation of the content analysis algorithm and its appeal to it lies in its consistency, i.e., in the mandatory stages of conscious learning, such as knowledge, understanding, application, analysis, and synthesis. The use of this method contributes not only to the processes of memorizing and reproducing facts, but also allows you to establish a connection with the previously received information, generalize, change, interpret and transform it.In this regard, the text is an invaluable source of factual material that provides options for the diverse use of language units at various levels. In Russian language classes, the most effective methods for semantic text analysis are the associative method and the content analysis method. Tasks based on associations arouse students ‘ cognitive interest, and content analysis contributes to better assimilation of texts of different speech styles, including texts on the specialtySemantic analysis of texts based on Bloom’s taxonomy in a non-linguistic university contributes to the development of critical and creative thinking of students by forming the ability to determine the value of an idea based on a critical and objective consideration of arguments

R.D. Darkembayeva, N. Ozekbayeva, F. Sametova · 0 citations
Open access Jul 2026

Human vs AI-Generated Texts in Language Learning: A Linguistic Comparison

The findings show that AI-generated texts exhibit greater lexical diversity and syntactic complexity; however, they often exhibit structural uniformity, overuse of cohesive devices, and limited pragmatic depth, and should not replace professionally designed educational materials.

V. Smaglii, T. Korolova, Svitlana Yukhymets et al. · 0 citations
Open access Aug 2026

Evaluating the Role of Artificial Intelligence in Supporting Usage - Based Approaches to Grammar

This study investigates grammatical trends in texts generated by artificial intelligence and human learners. The study puts to the test a fundamental principle of usage-based grammar: language is learned through repeated exposure to patterns. A direct comparison is conducted between AI-generated writings and language learners' essays. Quantitative approaches count words, sentences, and grammatical errors. Qualitative analysis detects trends in sentence structure and specific qualities such as past tense. Finding out if AI models adhere to usage-based grammar rules is the aim. Comparing the two groups' mistake types is another objective. The results show that whereas human writing varies, AI output is very constant. Almost no grammatical errors were found in AI articles, according to the study. Expected errors in human texts include omissions and overgeneralizations. The findings also demonstrate that AI makes greater use of components like the past tense and plurals. These studies demonstrate that the outcomes of usage-based learning are operationally replicated by AI. The results of training the model on massive amounts of data are consistent and precise. The ongoing process of language acquisition is reflected in human output. The study comes to the conclusion that AI is a powerful instrument for confirming frequency-based linguistic theory.but does not model the human cognitive journey. Future research should investigate different AI models and learner proficiency levels

Assis. lect. Batool Abdul-Mohsin Miri · 0 citations