Skip to content
Conference

ViCorpReviews: A Benchmark Dataset for Multi-Dimensional Sentiment and Hate Speech Detection in Vietnamese Workplace Context

Aug 2026 · International Conference on Multimedia Analysis and Pattern Recognition · pp. 192-197 · 0 citations · 16 references

Abstract

Online company review platforms have gained significant importance as they provide transparent insights into corporate culture and employee satisfaction. However, analyzing this feedback at scale remains challenging due to the linguistic complexity of workplace-specific discourse and the lack of high-quality, multi-dimensional annotated datasets, especially for low-resource languages like Vietnamese. This paper introduces ViCorpReviews, a novel benchmark dataset curated and annotated for two tasks: Aspect-based Sentiment Analysis (ABSA) and Hate Speech Detection (HSD). We implement a Human-in-the-loop annotation pipeline, leveraging the capabilities of Large Language Models (LLMs) to generate initial labels, which are subsequently refined and validated through a human verification process to ensure high-quality data. To establish robust baselines, we evaluate several well-known language models used for Vietnamese, specifically PhoBERT, XLM-RoBERTa, mT5, and ViT5. Our experimental results reveal the intricate patterns of workplace toxicity and sentiment, demonstrating that ABSA and HSD together provide complementary insights into the understanding of employee experience. The ViCorpReviews dataset serves as a foundational benchmark to foster future research in pattern recognition and natural language understanding within the corporate environment. Our dataset is released through this link: https://github.com/khanhlonguit/ViCorpReviews.

View source

Similar papers

Open access Aug 2026

Multi-Task BanglaBERT for Joint Sentiment and Fake News Detection in COVID-19 Discourse

The COVID-19 pandemic catalyzed an unprecedented surge of misinformation on social media, frequently intertwined with emotionally charged language. Understanding both the sentiment and truthfulness of this content is critical for public health monitoring and misinformation mitigation. However, Bangla—despite being a gl...

Arshadul Hoque · 0 citations
Conference

Leveraging LLM for Sentiment Detection in News Headlines

In this study, a longitudinal dataset of more than 23 million news headlines from 47 U.S.-based media outlets is used to investigate the use of large language models (LLMs) for sentiment detection. Recent developments in LLMs offer potential gains in contextual understanding, adaptability, and generalization, even thou...

Jae-Oong Yeom, YongKyung Oh · 0 citations
#large language models Open access Sep 2026

Evaluating fine-tuned, embedding-based, and zero-shot models for aspect-based sentiment analysis in South Slavic news

In the contemporary digital media landscape, the ability to automatically distill public opinion from a vast and continuous stream of information is highly important. Aspect-Based Sentiment Analysis (ABSA) offers this granular capability. In this work, we address a specific, industrially relevant formulation of this ta...

Nishan Chatterjee, B. Koloski, Antoine Doucet et al. · 0 citations
Open access Aug 2026

Indonesian Hate Speech Detection Across Diverse Domains Using Parameter-Efficient Fine-Tuning with IndoBERT and LoRA

The findings indicate that IndoBERT+LoRA provides a promising and resource-efficient approach for multi-domain Indonesian hate and abusive speech classification, while stricter leave-one-domain-out evaluation remains an important direction for future work.

Fergie Joanda Kaunang, Bhustomy Hakim, A. P. Thenata · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.