In digital pragmatics of CA on X, emoji use association with lexical/pragmatic category can be explained by a hybrid approach of computational, statistical, and pragmatic methods, reflecting the interaction among machine learning, linguistic/lexical features, contextual representation, and pragmatic communication.
Abstract
This study examines the performance of the state-of-the-art MARBERT model in identifying the lexical/pragmatic category associated with emoji use on X within a digital pragmatics approach (DPA). A net corpus of 15856 Colloquial Arabic (CA) posts containing emojis was collected from X using Python. The texts were tokenized and normalized into 4 lexical categories, namely noun_norm, verb_norm, adj_norm, and adverb_norm, and 2 pragmatic/structural categories, question_norm and exclamation_norm. MARBERT was finetuned and optimized to identify which category scores standard metrics more, hence associated with emoji use, while binary logistic regression was used to examine which category is statistically associated with emoji occurrence. Findings unveil that nouns dominate the corpus in normalized frequency (M = 0.675, SD = 0.161), followed by verbs (M = 0.083, SD = 0.100). However, verbs have the strongest influence of emoji use indicated by verb density (\b{eta} = 0.821, p = .001, 95% CI [0.332, 1.309]). The study concludes that in digital pragmatics of CA on X, emoji use association with lexical/pragmatic category can be explained by a hybrid approach of computational, statistical, and pragmatic methods, reflecting the interaction among machine learning, linguistic/lexical features, contextual representation, and pragmatic communication.
Talking openly about anxiety and depression (A&D) remains difficult for many people because of the stigma surrounding mental illness. Anonymous online platforms such as Reddit provide a space where users can express their thoughts and emotions more freely. This study considers how individuals linguistically construct and intensify emotional distress by examining (1) the adjectives used to express A&D, (2) the content-word collocates that co-occur with these adjectives, (3) the lexical features of these collocates, and (4) the emotional meanings conveyed through these collocational patterns. The dataset consisted of 1,440 Reddit posts (approximately 300,000 tokens) systematically sampled from the r/Anxiety and r/Depression subreddits between 2023 and 2025. An observational mixed-methods corpus linguistic approach was used to examine the data. Quantitative corpus linguistic analyses were carried out using AntConc, including frequency profiling and Mutual Information (MI) analysis, and were enhanced by qualitative concordance and Keyword-in-Context (KWIC) analysis to examine collocational patterns in context. The analysis shows a predominance of negatively valenced adjectives (e.g., anxious, depressed, hopeless, and suicidal), whose meanings are systematically intensified through their collocational environments. The collocates show distinctive lexical features. These include clinical nouns, linking and change-of-state verbs, and degree and frequency adverbs. These features construct varying levels of affective intensity and psychological distress. Emotional meaning is encoded in recurrent collocational patterns. Individual lexical items reveal only part of this meaning. This shows the value of collocational analysis for digital mental health research. The findings also possess practical implications. They may help improve the diagnostic sensitivity of automated digital mental health tools and foster more empathetic clinical communication.
Yu-Hang Yang, Kia-Lau Su, Handoko Handoko et al.· Jurnal Arbitrer· 0 citations
This study examines Transformer-based models'ability to learn emoji pragmatics in Arabic digital discourse (ADD), providing evidence from MARBERT's behavior with interpersonal pragmatic functions (IPFs). A corpus of 8,504 unique emoji-posts collected from Facebook via Python was used in the study. These posts were manually annotated, developed, and labeled for five IPFs: Politeness, Respect, Solidarity, Empathy, and Encouragement. A mixed-method approach was employed comprising statistical methods and interpretative analyses involving speech act theory, politeness theory, and rapport management theory. MARBERT was fine-tuned to model these context-dependent pragmatic functions. Findings demonstrate MARBERT's ability to learn these IPFs, achieving strong performance on unseen data, with an accuracy of 93%, a micro F1-score of 0.61, and a macro F1-score of 0.56, demonstrating its effectiveness in capturing interpersonal functions beyond conventional sentiment analysis. Function-level evaluation showed that Politeness and Respect were identified more accurately than Solidarity, reflecting differences in the explicitness and contextual dependence of IPFs. The study concludes that Transformer-based models learn patterns of face management and relational communication but remain challenged by highly implicit social meanings. It contributes a novel computational approach to modeling emoji pragmatics and advances the integration of interpersonal pragmatics with NLP for digital communication research.
The results show that semantic features consistently exhibit stronger relationships with readability than conventional lexical measures, indicating that semantic clarity has a greater impact on the readability of Indonesian academic writing than lexical sophistication alone.
Herpindo Herpindo, Sri Wulandari, Danang Agung Restu Aji et al.· Jurnal Sosioteknologi· 0 citations
The article substantiates the methodological value of the British National Corpus and the Corpus of Contemporary American English for the study of the verbal emoticon in contemporary English. The verbal emoticon is defined as a functional-discursive category: a verbal unit or construction that names, frames, intensifies, evaluates, or performs emotion in a specific context. The category includes lexical nominations, predicative patterns, evaluative exclamatives, interjective formulas, metaphorical expressions, and phraseological units when they function as markers of affective stance or interpersonal alignment. The study argues that BNC and COCA should not be used as interchangeable sources. BNC provides documented British evidence and enables comparison between written and spoken material, while COCA offers a large American corpus with genre-sensitive access to frequency, KWIC, collocation, and diachronic distribution. The article develops a corpus-based procedure that combines inventory building, concordance reading, collocational and constructional analysis, genre comparison, and contextual annotation. It also defines the interpretive limits of corpus evidence: frequency measures textual representation rather than psychological occurrence; collocation indicates recurrent co-selection but does not explain meaning without context; and corpus architecture determines the scope of possible conclusions. The final section outlines the relevance of corpus-grounded semantic profiling for large language models, especially for emotion recognition, stance detection, and affect-sensitive language generation.
N. Bober, R. Makhachashvili· Scientific Journal of Poloni...· 0 citations
The Türkiye Century Maarif Model, introduced by the Ministry of National Education in 2024, represents a major curriculum reform. Although previous research has examined teachers, administrators, and teacher candidates mainly through surveys and interviews, spontaneous digital reactions remain underexplored. This study investigates discourse surrounding the Maarif Model on X. Posts containing five reform-related hashtags were collected through the X API and analyzed in Python using a KDD-informed text-mining framework. The analysis combined BERT-based sentiment classification, word-frequency analysis, and Latent Dirichlet Allocation topic modelling. Data preparation included duplicate removal, relevance screening, and bot/noise filtering, while automated sentiment labels were interpreted cautiously because no independently coded gold-standard validation set was available. Positive posts constituted the largest sentiment category across all hashtag-based subsets, suggesting substantial symbolic support for the reform. However, negatively oriented posts contained more prominent references to projects, targets, administrative expectations, and implementation-related concerns. Topic patterns also linked the reform discourse to curriculum, teacher roles, student-centered approaches, institutional coordination, and local implementation. Broader hashtags were associated with relatively more critical or contested expressions, whereas institutionally aligned hashtags contained stronger positive discourse. These findings suggest that social media can provide an early source of feedback on large-scale educational reforms, although it should not be treated as representative public opinion. The study illustrates how computational analysis of digital discourse can complement traditional policy-evaluation methods.
Ahmet Başal, M. Kayakuş, Cem Oktay Guzeller et al.· Participatory Educational Re...· 0 citations
This study aims to identify, classify, and analyze collocational patterns in Surah Yusuf and to examine their semantic roles. The method employed is a descriptive qualitative approach using manual concordance analysis, in which data were extracted through close, word by word reading of the Arabic Qur'anic Text of Surah Yusuf. Identified collocations were verified using the Quranic Arabic Corpus to confirm their morphological and syntactic structure, and cross checked against Arabic collocation dictionaries to ensure linguistic validity. Data were analyzed based on collocation types (grammatical and lexical) and the semantic patterns of the accompanying prepositions. The findings reveal 36 collocational patterns, dominated by grammatical verb collocations (26 patterns), particularly those using fiʿl māḍī (past tense verbs). Prepositions such as عَنْ (for avoidance), إِلَى (for direction), and عَلَى (for authority) form consistent semantic constructions that enhance narrative nuances. In-depth analysis of key collocations, such as رَاوَدَ عَنْ and صَبْرٌ جَمِيْلٌ, shows that collocational choices do not merely serve descriptive functions but actively shape character development, dramatic tension, and core theological concepts. The study concludes that collocational analysis is an essential tool for uncovering complex layers of meaning in the Qur'an. These findings make a significant contribution to Qur'anic linguistics and exegesis by offering a systematic and replicable model for analyzing other surahs. The research also has practical implications for translation, the development of Arabic language learning materials, and digital humanities.