Skip to content
Preprint

From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music

Aug 2026 · 0 citations · 26 references
Computer Science

TL;DR

This work pair 200 real-world Udio prompts with their generated audio and free-form descriptions collected from English- and Korean-speaking listeners, and contributes a human-derived taxonomy of musical prompting vocabulary grounded in real user data, finding a consistent structural asymmetry.

Abstract

Text-to-music (TTM) generation systems allow users to create music through natural language prompts, yet it is unclear whether the descriptive language used to prompt aligns with descriptive language used to summarize or describe heard music. We pair 200 real-world Udio prompts with their generated audio and free-form descriptions collected from English- (n = 70) and Korean-speaking (n = 78) listeners, and contribute a human-derived taxonomy of musical prompting vocabulary grounded in real user data. Using this framework, alongside word- and vector-level analyses, we find a consistent structural asymmetry: prompts are dominated by Genre and Story/Narrative language. Genre terms propagate most reliably from prompt to perception, while narrative-heavy prompts are the strongest predictor of semantic misalignment. A preliminary cross-cultural comparison further suggests that description profiles vary across listener populations along narrative, functional, and affective dimensions, raising questions about whether current TTM systems, trained on aggregated English-centric corpora, can accommodate the full diversity of how people naturally express musical ideas.

View source

Similar papers

Preprint Aug 2026

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct"naturalness"into a...

Oluwanifemi Bamgbose, Simon Rosen, J. Shah et al. · 0 citations
Review Open access Aug 2026

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Ashkan Goudarzi, Aylar Naderi Zonouz · 0 citations
#artificial intelligence Preprint Jul 2026

Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study

Audio-language models increasingly generate confident music descriptions that are unsupported by the input audio. We present, to our knowledge, the first music-specific, layer-wise, multi-paradigm empirical study of hallucination in audio-language models and formulate it as a hierarchical perceptual grounding failure a...

Yu Liu, Jia-Hui Liu, Zhi-Lin Liu et al. · 0 citations
Preprint Aug 2026

What Are You Listening to? Temporal Music Grounding for Audio-to-Text Large Language Models

Large audio-language models can produce fluent and musically plausible responses, yet it often remains unclear whether those responses are grounded in the audio input. We introduce temporal music grounding, a task in which a model returns one or more time spans corresponding to a queried musical note, event, or pattern...

Kun Fang, Ziyu Wang, Ichiro Fujinaga · 0 citations
Open access Aug 2026

Translingual Resonance in Popular Music: A Comparative Study of Analogue, Platform-Mediated, and AI-Assisted Circulation

Music can support affective engagement across linguistic boundaries, but the mechanisms and infrastructures of that engagement remain dispersed across sociolinguistics, translation studies, music cognition, and platform research. This conceptual article compares three regimes of cross-lingual musical circulation: analo...

Smriti Tripathi · 0 citations
#small language model Open access Sep 2026

Human versus machine

Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost, genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for...

Afra Sturm, Valentin Unger, Fabian Grünig · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.