Speech Editor: an interactive speech editing software with disentangled representation learning
Abstract
Speech editing is gradually making its way into everyday applications. This paper presents Speech Editor, a practical software system built upon SpeechTripleNet for disentangled speech representation learning. The system enables real-time end-to-end speech editing, including local modification of pitch and energy as well as speaker conversion without requiring external labels. The software integrates a lightweight backend with a user-friendly frontend, allowing non-specialist users to interactively edit speech through a browser interface. Experimental results demonstrate that the model correctly disentangles the speech representations and supports speech editing. The proposed system offers high usability, portability, and fidelity, making it suitable for applications in podcasts, dubbing, and privacy protection.