From direct scoring to cue integration: big five prediction from spoken responses to a picture-based self-projection task
Abstract
Open-ended language tasks are increasingly used in personality assessment, but their psychometric value largely depends on two factors: the type of evidence a task elicits and how that evidence is represented before scoring. In this study, we examined five spoken responses to a picture-based self-projection task collected from 374 participants during in-person workplace training sessions. Participants first completed the Big Five Inventory-2 questionnaire; the spoken task was administered approximately 1 day later, and responses were transcribed using a Whisper-based automatic speech recognition pipeline. We compared four approaches applied to the same transcribed corpus: Linguistic Inquiry and Word Count (a lexical dictionary method), contextual semantic embeddings, holistic large language model (LLM) appraisal (holistic appraisal route; HAR), and hierarchical cue integration route (HCIR). The results were trait-specific: extraversion was the most recoverable domain across methods, whereas Openness was the least recoverable. HAR preserved rank-order information, particularly for Extraversion, but exhibited poor calibration, evidenced by negative mean R ² values and systematic directional bias in extreme score bands. In the retained comparison, HCIR-Grounded showed the most favorable observed association–calibration profile among the evaluated approaches (mean | r | = 0.264, mean R 2 = 0.073), with notably higher observed correlations for Agreeableness, Conscientiousness, and Neuroticism. We refrain from attributing this pattern solely to cue integration, however, as HCIR included a supervised calibration stage that HAR lacked; the comparison therefore reflects the joint contribution of cue organization and supervised calibration. By explicitly organizing cue provenance, self-relevance, quality indicators, and trait-specific evidence before supervised calibration, HCIR supports a measurement-oriented interpretation of LLM-assisted personality inference from spoken responses that combine self-projection, preference, and normative evaluation.