Skip to content
Open access

Research on Deep Learning-Based Methods for Emotional Expression Recognition and Style Analysis in Piano Performance

Aug 2026 · Advanced Electromagnetics · 0 citations · 7 references

Abstract

Emotional expression recognition in piano performance requires fine-grained modeling of audio timbre, dynamic touch, rhythm fluctuation, and performance style. Existing methods based on single audio features or shallow statistics often fail to capture subtle emotional transitions and stylistic differences. This study proposes a deep-learning-based framework integrating multimodal feature fusion, temporal modeling, attention mechanisms, and emotion-style collaborative learning. Piano audio signals are converted into time-frequency representations through short-time Fourier transform and Mel-scale mapping, while MIDI control information encodes velocity, duration, and beat-offset features. Audio spectrogram features and MIDI control vectors are fused to construct a unified representation. A convolutional network extracts local time-frequency patterns related to accents and touch variations, and a temporal modeling module captures emotional evolution across performance sequences. An energy- and velocity-guided attention mechanism enhances key musical segments, including climaxes, strong beats, and ornaments. In addition, a style representation branch models rhythmic stability, dynamic fluctuation, and tempo evolution, while cross-task consistency constraints couple emotion recognition and style analysis. Experiments on paired audio–MIDI performance data show that the proposed method achieves 0.91 accuracy and 0.90 F1 score, improving by approximately 9% and 10% over unimodal baselines. The framework provides an audio signal processing and multimodal sequence-analysis method for intelligent music-performance evaluation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.