Skip to content
Open access

The Corpus of English–Tagalog Code-switching: an integrative corpus of linguistic contact effects

Sep 2026 · Scientific Data · 0 citations

Abstract

The Corpus of English–Tagalog Code-switching (CEnTaCS) is a new corpus designed to document naturalistic English–Tagalog bilingual speech. Locally referred to as Taglish, it is a widely-used yet understudied contact variety in the Philippines. Recorded in 2025 in Metro Manila, the corpus comprises three interrelated datasets: individual story retelling recordings, cognitive control data based on the Arrow Flanker task, and sociolinguistic questionnaire data. CEnTaCS offers fine-grained data on code-switching across clausal, lexical, and morphemic levels, with particular attention to intra-word code-switching, a salient but still insufficiently documented feature of Taglish. Beyond code-switching, the corpus captures a broad range of language contact phenomena, including borrowing, calquing, convergence, and interference. By making these data systematically available, the corpus supports an integrated account of mixed-language practices that brings together structural, cognitive, and sociolinguistic perspectives. CEnTaCS thus enables researchers to investigate how contact-induced structures are constrained linguistically, how they relate to bilingual processing, and how they index social meaning and identity in a community where bilingualism is deeply embedded in everyday life.

Read PDF

Similar papers

Open access Aug 2026

Exploring Mixed Copies: Evidence From English-Estonian Bilingual Speech

Recently, contact linguistics has become increasingly interested in multiword units. At the same time, the code-copying framework (CCF) includes the notion of mixed copies (MCs) that are in-between global copies (‘borrowing’) and selective copies (‘structural change’) and illustrate the transition between the lexic...

A. Verschik, H. Kask · 0 citations
Aug 2026

Spanish–Guaraní Code-Switching in Social Media: A Dependency Parsing Study

This study examines how Spanish–Guaraní bilinguals structure code-switching (CS) on Twitter/X and investigates what these patterns reveal about competing models of bilingual grammar. Specifically, it explores language dominance, switch direction, syntactic position, switch span, and register effects in Paraguayan b...

Olga Kellert · 0 citations
Open access Sep 2026

English-Irish codeswitching as a stylistic resource in stancetaking: a case study of the queer bilingual podcast Gaylinn

Abstract This article examines how codeswitching (CS) from English into Irish (Gaeilge) functions as a resource for stancetaking in contemporary queer media. Drawing on a corpus of spontaneous speech from the queer bilingual podcast Gaylinn, the study integrates variationist, interactional, and new-speaker-oriented the...

D. Mooney · 0 citations
Open access Sep 2026

Exploring the MetaLing Corpus : spelling variation and discourse in Early Modern English texts on language

This article explores the analytical potential of the MetaLing Corpus , a collection of Early Modern English texts on language dating from 1500 to 1700. Combining corpus-driven and corpus-based approaches, it demonstrates both the opportunities and the challenges of working with a corpus currently under developme...

A. Andreani, Vahid Asadi · 1 citation
Open access Sep 2026

பெருமாள் முருகன் நாவல்களில் கொங்கு வட்டார மொழியும் உள்ளூர் அறிவு மரபுகளும்: Corpus Linguistics மற்றும் Digital Humanities அடிப்படையிலான ஆய்வு

This article examines the Kongu regional linguistic elements found in these novels and the local knowledge traditions they reveal and the need for and possibilities of documenting regional dialect literary works through digital humanities tools using Corpus Linguistics and Digital Humanities methodologies.

முனைவர் இல. பூவலிங்கம் · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.