A reinforcement learning-based optimization method for personalized teaching models in college piano courses
Abstract
Traditional university piano teaching has long been limited by static rule interventions and single-dimensional outcome evaluations, making it difficult to adapt to the nonlinear and dynamic development needs of individual cognitive levels and performance skills. Exploring intelligent auxiliary paths with real-time feedback mechanisms has become crucial for the engineering of music education. To this end, this paper proposes a personalized teaching model optimization method based on a deep deterministic policy gradient algorithm. This method reduces the dimensionality of temporal acoustic data and reconstructs it into a 15 -dimensional continuous Markov state space that integrates pitch shift, tempo rhythm, and dynamic gradient. Simultaneously, it defines a 3- dimensional continuous action space encompassing repertoire difficulty, slice practice length, and intervention frequency, and establishes a multi-objective reward function integrating cognitive load penalty terms to guide the smooth convergence of teaching decisions. Simulation and quantitative test results show that the proposed optimization algorithm reduces the policy convergence period to 640 rounds, and the overall cumulative expected return score remains stable at 92.8. In the evaluation of teaching application effectiveness, this model improves the average skill mastery of the test sample group to 88 points and effectively improves the long-tail distribution characteristics, with a reduction of over 60% in the variance of group learning differences. This study theoretically verifies the feasibility of the continuous action space decision model in processing complex unstructured educational data, providing underlying architectural support and methodological foundation for the adaptive education paradigm upgrade and large-scale intelligent deployment of advanced art practical skills courses.