Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
DSSM-CRF, an audio-only architecture that explicitly separates cross-speaker contextual influence and within-speaker emotion evolution, is proposed and matched controls demonstrate complementary gains from speaker-wise factorization and CRF modeling.