A Distributed Learning Approach for Controlled Grammar Transfer using Structure–Timbre Disentangled Audio Synthesis
Abstract
This research article introduces a probabilistic framework for controlled grammar transfer in audio synthesis, tackling the issue of blending audio styles in a way that is both interpretable and disentangled. Current methods either depend on low-level feature matching or use neural models that mix structural and timbral representations, which prevents precise control over what is transferred. We suggest a system based on a Hidden Markov Model within a distributed client-server architecture, where transition matrices represent temporal grammar and Gaussian emission parameters capture timbre. Grammar transfer is achieved through convex interpolation of aligned transition matrices, controlled by a scalar variable, while audio is generated by selecting frames solely from the source instrument’s state-conditioned banks, inherently preserving timbre. Experiments on the IRMAS dataset demonstrate that the proposed framework achieves monotonic grammar interpolation, reducing the Jensen–Shannon divergence between generated and donor transition matrices from approximately 0.13 for the strongest baseline to below 0.01, corresponding to about a 95% reduction, while maintaining comparable timbre preservation and signal quality.