Skip to content

Building Trustworthy Mental Health Benchmarks on Bluesky: A Validation-Aware Weak-Supervision Framework

Sep 2026 · 0 citations · 15 references
Computer Science

TL;DR

These findings show that decentralized social media can support reproducible mental health benchmarking, but only when system design, label provenance, validation strategy, and deployment constraints are evaluated together.

Abstract

Decentralized social media platforms create new opportunities and challenges for computational mental health research because data access, moderation, labeling, and deployment responsibilities are distributed across multiple technical and governance layers. This paper presents a validation-aware weak-supervision system for constructing and evaluating suicidal ideation (SI) and broader mental health (MH) disclosure benchmarks on Bluesky, a decentralized social media platform built on the AT Protocol. The system integrates public firehose collection, task-specific lexicon filtering, Llama-3-8B-assisted binary annotation, human-adjudicated validation subsets, and transformer-based model benchmarking. Using this pipeline, we construct two task-specific corpora containing 8,346 SI-labeled posts and 9,988 MH-labeled posts. The evaluation shows that model performance depends strongly on both task definition and validation protocol. BERT+LSTM achieves the highest SI stratified cross-validation F1-score, RoBERTa achieves the strongest SI holdout F1-score, and DistilRoBERTa achieves the best MH cross-validation F1-score. Human validation reveals different weak-label failure modes across tasks, with SI labels dominated by false negatives and MH labels dominated by false positives. These findings show that decentralized social media can support reproducible mental health benchmarking, but only when system design, label provenance, validation strategy, and deployment constraints are evaluated together.

View source

Similar papers

Open access 2026

CueSupport-MH: An Explainable Dataset and Interpretable Framework for Mental Health Risk Detection and Response Safety Evaluation

CueSupport-MH is introduced, a 6,400-instance benchmark constructed from public Reddit conversations and annotated for four complementary dimensions: risk label, risk evidence span, protective cue, and response-safety label, and a lightweight interpretable framework for joint risk detection and response-safety classifi...

F. Alotaibi, Msvpj Sathvik, Mohammed Younus et al. · 0 citations
Open access Sep 2026

Modeling Co-Existing Mental Health Risks in Social Media via Multi-Label Learning and LLM

Mental health risk detection from user-generated social media text has become increasingly important as psychiatric conditions such as depression, anxiety, and suicidal ideation continue to rise. However, most existing computational studies operationalize the problem using single-label datasets, implicitly assuming mut...

Z. Özer, Emre Kenger, Görkem Gözükara et al. · 0 citations
Review Sep 2026

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and fou...

G. Guérard, Sonia Djebali · 0 citations
Conference Sep 2026

A Causally Informed CNN-Ensemble Fusion Framework for Reddit-Based Mental Health Using Textual, Temporal, and Metadata Features

Social media offers large-scale textual data on social indicators of distress, support-seeking and disclosure of psychological crises. Nevertheless, most mental health classification systems remain prediction focused and offer little information about how textual, temporal, and metadata signals interact to predict ment...

U. Nazir, Nur Shazwani Kamarudin, M. A. Majid et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.