Inferring Big-Five Personality Traits via Heterogeneous User-Post Graph Representation Learning
Abstract
Automated inference of Big Five personality traits from social media text is a growing research area with applications in personalized recommendation and human-computer interaction. Most existing approaches treat user posts as independent textual units, overlooking structural relationships between users and their content. This paper proposes a Heterogeneous User-Post Graph Representation Learning framework that models users and posts as distinct node types connected through authorship and semantic-similarity relations, processed by a heterogeneous Graph Neural Network (HeteroGNN) built on contextual sentence embeddings and FAISS-based nearest neighbor search. Evaluated on the large-scale Pandora dataset against a transformer-based baseline (RoBERTa) and traditional machine learning baselines (Decision Tree, Random Forest), RoBERTa achieved the lowest average RMSE (0.2550), while HeteroGNN (0.2864) outperformed both Random Forest (0.2885) and Decision Tree (0.3231) across all five dimensions (p < 0.001). A controlled ablation found no benefit from heterogeneous modeling when user nodes lacked informative content, but a significant, consistent benefit (1.5–2.1% RMSE reduction) once user nodes carried genuine content from their associated text. These findings indicate that heterogeneous relation-type modeling is beneficial specifically when differentiated node types independently carry informative signal, and that the remaining gap to transformer-based methods stems from a lack of task-adaptive, jointly-attended representations.