Jun 2026· Training Language and Culture· Vol 10, pp. 20-35· 0 citations
Abstract
In recent years, corpus linguistics has provided new opportunities for studying language contact using quantitative and data-driven techniques. For the Kazakh language, situated in a multilingual setting involving Kazakh, Russian, and English, lexical interference represents an important sociolinguistic issue. This study investigates lexical competition from quantitative and functional standpoints, with particular attention to frequency, distribution, and variation in Kazakh-language mass media under Kazakh-Russian language contact. Using corpus methods, the study shows that large collections of texts make it possible to identify borrowed lexemes and competing equivalent forms. Borrowed items may coexist with native equivalents, producing measurable differences in frequency and stylistic preference. In some instances, foreign lexemes are more frequent than standardised national alternatives in the analysed corpus, although established Kazakh norms may continue to coexist with such borrowings. To address this issue, the study introduces the Lexical Competition Coefficient (LCC), a descriptive numerical indicator that measures the relative frequency of competing lexical equivalents in corpus data. When the Kazakh item predominates, the LCC exceeds one (k > 1), indicating that the Kazakh variant is more frequent in the corpus. By contrast, a coefficient below one (k < 1) indicates that the borrowed variant is more frequent. The findings indicate that lexical competition in Kazakh-language media is uneven across lexical pairs, with some borrowed forms substantially outnumbering Kazakh equivalents while other Kazakh variants remain dominant in corpus usage. The study combines scholarship on language contact and quantitative corpus analysis and proposes a systematic procedure for assessing lexical competition in corpus data.
The internationalization of higher education and scholarly publishing has made English an increasingly multilingual medium of academic communication. This study examines similarities and differences between native and non-native academic English through a computational corpus-based framework. The study aims to identify variation in lexical choices, lexical bundles, grammatical patterns, collocations, and academic stance. A comparative corpus design is proposed using matched academic texts produced by native and non-native English writers. Computational procedures include normalized frequency analysis, keyword analysis, n-gram and lexical-bundle extraction, collocation analysis, part-of-speech profiling, and stance-marker analysis. Recent corpus research indicates that lexical bundles and stance resources provide measurable evidence of variation in L2 academic writing (Chen, 2025; Siu et al., 2024), while World Englishes scholarship increasingly challenges the assumption that native-speaker English should function as the exclusive model for academic writing (Du & Liu, 2025). The study therefore approaches native and non-native academic English as analytically comparable but potentially diverse forms of scholarly communication rather than as superior and deficient varieties. The proposed analysis is expected to identify both shared academic conventions and systematic differences in phraseological, lexical, grammatical, and interpersonal choices. The study contributes to corpus linguistics and English for Academic Purposes by providing a computationally informed framework for examining academic language variation and by offering implications for corpus-based academic writing pedagogy.
N. Akhter, Fatima Khan, Hammad Malik et al.· Journal of Arts and Linguist...· 0 citations
The transition of Uzbekistan to a market economy and its integration into global economic networks have driven an intensive influx of economic terminology into the Uzbek language, predominantly from two donor languages — Russian and English. This article examines the lexical-semantic features that accompany the assimilation of such borrowed economic terms into the recipient language. Drawing on a corpus of approximately 450 economic term units compiled from lexicographic sources, specialized textbooks, normative-legal texts and contemporary economic media, the study applies descriptive-comparative, componential and definitional-contextual analysis to trace how meaning is restructured during borrowing. The findings distinguish four principal source strata — direct English borrowings, Russian-mediated internationalisms, direct Russian borrowings and calques built from native material — and identify five recurrent semantic-adaptation mechanisms: full semantic transfer, semantic narrowing, semantic widening, semantic shift and loan translation. Particular attention is given to synonymic doublets and triplets in which native, Russian and English designations compete and undergo stylistic and functional redistribution, and to the polysemy and homonymy generated by multi-channel borrowing. The results demonstrate that economic borrowing in Uzbek is not a mechanical transfer of labels but a systematic semantic reorganization governed by donor-language chronology, register and the word-formation resources of the recipient language.
Muhayyo Qutliyeva· Ijtimoiy-gumanitar sohada il...· 0 citations
Recently, contact linguistics has become increasingly interested in multiword units. At the same time, the code-copying framework (CCF) includes the notion of mixed copies (MCs) that are in-between global copies (‘borrowing’) and selective copies (‘structural change’) and illustrate the transition between the lexicon and grammar. The research question is: What types of MCs occur in English-Estonian bilingual speech?
The data were transcribed, and MCs identified, annotated, and classified according to their structure. English items were searched for in Estonian dictionaries to establish their Estonian equivalents or conventionalization of such items. The frequencies of MCs and their Estonian equivalents were also searched on Google to determine whether the MCs occur outside the corpus.
Three datasets were analysed: written texts from 44 blogs (385,124 tokens), spoken data from 10 vlogs (117,555 tokens), and 8 podcasts (77,277 tokens). Quantitative analyses of the various MC types were conducted, followed by a qualitative analysis of representative examples.
Compound nouns constitute the majority of MCs, followed by idioms, phrasal compounds, and a small number of compound verbs. No frame-changing MCs (i.e., MCs resulting in grammatical change) were attested. Since compound nouns and analytic verbs occur in both languages, structural similarity may be a facilitating factor in copying.
The notion of MCs is not widely used. Research typically focuses on particular types of items (e.g., compound nouns or verbs); here, however, the question is reversed: which types of items yield MCs?
It was established that the proportion of MCs in the data is comparable to that of selective copies. Within MCs, the globally copied element renders the remaining part more specific, highlighting the importance of meaning in contact-induced language change. MCs are also present on the Estonian internet and, in some cases, outnumber their Estonian equivalents, if such equivalents exist.
A. Verschik, H. Kask· International Journal of Bil...· 0 citations
Conversion is a fertile resource of word-formation in English. Its degree of productivity is often reflected in the high incidence of converted lexemes in language resources, for instance dictionaries and corpora. With the main aim of surveying the alleged relationship between morphological productivity and frequency, this study examines English verbs derived by denominal conversion between 1990 and 2019. The analysis of a dataset compiled from the Oxford English Dictionary (OED) and the Corpus of Contemporary American English (COCA) reveals a concentration of converted verbs within the lowest frequency bands. The scarcity of converted verbs in mid-to-high frequency bands suggests that, while morphologically viable, they rarely achieve broad lexical entrenchment. This underscores the often ephemeral nature of the output of highly productive word-formation processes, with many formations remaining marginal or specialized in use, contingent upon their semantic categorization. It is noteworthy that, while the OED documents a significant number of converted verbs, the COCA data indicates that many of these remain peripheral in discourse. The prominence of converted verbs in the OED’s lowest bands suggests their recognition without widespread usage, consistent with prior observations that conversion frequently yields nonce formations and contextdependent lexical items. The semantic distribution of converted verbs further elucidates their productivity and frequency patterns. Instrumental verbs dominate, reflecting technological influence, while performative verbs also show remarkable representation. Overall, the findings confirm that denominal conversion is a productive but predominantly low-frequency process, shaped by morphosyntactic and semantic constraints.
Jesús Fernández-Domínguez· Arbeiten aus Anglistik und A...· 0 citations
This study employs the comparative-inductive method prevalent in research on cross-linguistic transfer within multilingual production. Taking Chinese, English and Japanese writing samples collected from nine student participants as research data, it conducts cross-text comparisons to investigate lexical errors in tri-lingual Japanese writing and clarify core issues including the sources and frequency distribution of first-language (L1) and second-language (L2) transfer. The results demonstrate that content words yield a higher overall error rate than function words. Statistically significant differences exist in error rates among subcategories of content words (namely nouns, verbs, adjectives and adverbs), whereas no significant differences are detected within the category of function words, which consist of auxiliary verbs, particles and other function-word items. From an individual learner perspective, both L1 and L2 transfer exhibit considerable individual variability in the use of either content or function words, with far more prominent divergence observed in content word usage. On the whole, content words involve a higher proportion of cross-linguistic transfer than function words, and instances of L1 transfer appear substantially more fre-quently than those of L2 transfer
You Ma, Hui Shi· Social Sciences and Humaniti...· 0 citations
This study examines how Spanish–Guaraní bilinguals structure code-switching (CS) on Twitter/X and investigates what these patterns reveal about competing models of bilingual grammar. Specifically, it explores language dominance, switch direction, syntactic position, switch span, and register effects in Paraguayan bilingual discourse.
The study adopts a corpus-based and dependency-analytic approach using Universal Dependencies (UD) annotation to examine CS patterns in a Spanish–Guaraní social media corpus.
The analysis is based on an extended UD-annotated Spanish–Guaraní Twitter/X corpus. Switch points were examined according to direction, dependency relations (DEPREL), part-of-speech categories, span length, and register (formal vs. informal tweets). Quantitative analyses were complemented by qualitative examination of representative examples.
Results show that Guaraní, including the variant Jopará, functions as the dominant matrix language overall, especially in formal tweets. Switches into Spanish occur disproportionately at core grammatical positions such as subjects, objects, and obliques, whereas switches into Guaraní frequently occur at clause-edge and discourse-related positions. Switch spans are generally short but vary systematically according to syntactic environment, with clause-structuring elements licensing longer spans than argument-level switches.
The study introduces switch span as a measure of structural scope in CS and combines dependency-based syntactic analysis with register-sensitive investigation in a low-resource bilingual setting.
The findings support Matrix-Language Frame predictions in formal registers while demonstrating the importance of hierarchical and scope-based accounts of bilingual grammar. More broadly, the study shows how dependency annotation can provide fine-grained insights into CS in Indigenous and mixed-language repertoires.
Olga Kellert· International Journal of Bil...· 0 citations