Linguistic Stereotyping, Gendered Representation, and the Shifting Portrayal of Women in Hindi Cinema: A Sociolinguistic Corpus Study
Hindi popular cinema in its studio era as well as in its OTT era has been an enduring medium through which speaker types have been socially encoded. The paper uses a bounded corpus of 2025 Hindi language feature films (40 theatrically released and two long-form web series) to explore how the language and paralinguistic decisions create and produce stereotypes of regions, religion, class, occupation, and gender, and how these decisions have changed in the OTT era. The unit of analysis is the character-scene, which is an individual character in an identifiable scene; utterances, accent features, lexical selection, address form and code-mixing are coded as embedded units. Operationalization of six stereotype categories is done by explicit indicator sets and not by impressionistic labels. Gendered language is distinguished from sexist language as the former is grammatical or referential, the latter evaluative and derogatory, placing women in an inferior, subordinate, or consumable position. Dialogue in Hindi–Urdu language is consistently romanised, and each cited example includes a romanised original, an English translation, identification of the speaker, and context for the scene. The paper shows that, while the stereotype grammar of Bihari-as-comic, Christian-as-westernized, English-speaking-woman-as-morally-suspect has been somewhat replaced, it has not been eliminated from the OTT corpus and that gendered-language markings continue to be used even though overt sexist framings are absent.