VIBE is introduced, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space and its core contribution is a measurement contract, which motivates entity-centered affective profiling as a documented practice.
Abstract
Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable. Existing work captures parts of this space through sentiment, favorability, and emotion benchmarks, but none combines target-directed VAD attribution, an explicit scorer contract, and a passport reporting format. We introduce VIBE, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space. Its core contribution is a measurement contract: VIBE separates generation from external scoring, distinguishes scalar favorability, response-level VAD, and target-directed VAD, and reports profiles through an Affective Passport. Three empirical layers support the contract. H1 shows scalar favorability does not subsume arousal and dominance: valence findings are cross-validated (rV = 0.944 judge-human, rV = 0.954 inter-scorer); arousal and dominance are single-scorer directional estimates, not point-precise, consistent with known inter-annotator difficulty on these axes (rA = 0.495, rD = 0.702 among human annotators). H2 shows whole-response and target-directed VAD are different contracts: the same text can carry one affective tone overall while representing the named target differently. H3 is a protocol-drift diagnostic: elicitation conditions shift profiles, motivating context metadata in every affective report. These results motivate entity-centered affective profiling as a documented practice: profiles should be released with scorer identity, coverage, protocol, and interpretation limits.
People often assume that emotionally intense relationship experiences "leak" into language: more elaborated stories should sound more emotional. We tested this expressive-congruence assumption in 351,734 de-identified relationship narratives using an interpretable index of narrative complexity (Level of Complexity; LoC...
GenAI and LLMs have the potential to enhance the affective intensity of communication without necessarily compromising its informational quality, however, as human communication becomes increasingly and inevitably AI-mediated, current LLMs may shift the affective register, narrow the emotional spectrum, and introduce t...
Samuel B. Mazzone, J. Harlan, K. M. Sajjadul Islam et al.· ACM Transactions on Social C...· 0 citations
Results show systematic variation across domains: technical and methodological areas such as deep learning and natural language processing exhibit gain-salient framing, while safety-critical topics such as deepfakes and facial recognition show strongly loss-salient profiles.
O. Topal, Inna Novalija, Joao Pita Costa et al.· Applied Informatics· 0 citations
Sentiment analysis, crucial in fields like human computer interaction and mental health, has advanced through multimodal fusion and large-scale language models. People exhibit significant variability in emotional expressions due to behavioral differences, yet existing methods often apply generic models across populatio...
Wenjuan Gong, Jiarui Li, Tingbo Shi et al.· ACM Transactions on Multimed...· 0 citations
Empathetic response generation requires models to decide not only what to say, but also how to respond to the previous speaker's affective situation, so multi-annotator emoji distributions are used as weak affective--attitudinal evidence to induce a latent control space that operationally approximates listener stance.
Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single i...
Zheng Wu, Yibo Luo, Pu Zhang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.