Skip to content
Open access

A Computational Corpus Study of English Varieties: Comparing Native and Non-Native Academic Discourse

Jul 2026 · Journal of Arts and Linguistics Studies · 0 citations · 6 references

Abstract

The internationalization of higher education and scholarly publishing has made English an increasingly multilingual medium of academic communication. This study examines similarities and differences between native and non-native academic English through a computational corpus-based framework. The study aims to identify variation in lexical choices, lexical bundles, grammatical patterns, collocations, and academic stance. A comparative corpus design is proposed using matched academic texts produced by native and non-native English writers. Computational procedures include normalized frequency analysis, keyword analysis, n-gram and lexical-bundle extraction, collocation analysis, part-of-speech profiling, and stance-marker analysis. Recent corpus research indicates that lexical bundles and stance resources provide measurable evidence of variation in L2 academic writing (Chen, 2025; Siu et al., 2024), while World Englishes scholarship increasingly challenges the assumption that native-speaker English should function as the exclusive model for academic writing (Du & Liu, 2025). The study therefore approaches native and non-native academic English as analytically comparable but potentially diverse forms of scholarly communication rather than as superior and deficient varieties. The proposed analysis is expected to identify both shared academic conventions and systematic differences in phraseological, lexical, grammatical, and interpersonal choices. The study contributes to corpus linguistics and English for Academic Purposes by providing a computationally informed framework for examining academic language variation and by offering implications for corpus-based academic writing pedagogy.

Read PDF