Skip to content
Preprint

Steering dense music retrieval with open-vocabulary concept discovery

Aug 2026 · 0 citations · 37 references
Computer Science Engineering

TL;DR

A lightweight, training-free method that recovers a sparse set of audio features whose decoded representation reconstructs the target concept while remaining consistent with audio-space geometry, enabling more precise concept amplification and suppression with reduced drift on preservation metrics.

Abstract

Controllable music retrieval lets users find music that is, for example, more ambient, less distorted, or without guitar while preserving the other semantic content of an original seed query. Sparse autoencoders (SAEs) are a promising interface for this kind of concept-level control, but a key problem remains: given a free-form text concept, which sparse features should be edited? In shared multimodal embedding spaces, standard attribution methods often select neurons that match the concept's wording but not the audio examples that express it. This leads to weak or unstable edits: relevant features are missed when concepts are distributed across neurons, while others are selected due to text alignment rather than audio-side structure. We address this with a lightweight, training-free method that recovers a sparse set of audio features whose decoded representation reconstructs the target concept while remaining consistent with audio-space geometry. This reframes concept attribution as a sparse inversion problem rather than a text-side neuron-ranking heuristic. The method requires neither paired audio-text supervision nor SAE retraining. We evaluate this approach in steerable music retrieval and show that the recovered supports align more closely with concept-bearing audio examples and achieve a stronger trade-off between edit strength and preservation than alignment baselines, enabling more precise concept amplification and suppression with reduced drift on preservation metrics.

View source

Similar papers

#machine learning Preprint Sep 2026

If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization

Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sources available to this task can differ in temporal-boundary reliability: on OV-AVEBench, our configured visual teacher gives more reliable boundary cues than the configured...

Yi Xu, Cheng Chen, Wen-Zhuo Lei · 0 citations
Book Open access Aug 2026

LSAR: Sparse Lexical Representation Learning for Efficient and Interpretable Audio Retrieval

As Multimodal Large Language Models (MLLMs) expand the scope of retrieval-augmented generation, recommendation, and multimedia search, audio retrieval is expected to become a dependable retrieval component. Yet existing systems struggle to reconcile lexical precision, non-verbal acoustic evidence, and efficient, transp...

Haoyu Li, Yuzhe Bai, Li Niu · 0 citations
Preprint Aug 2026

Learning Sample-wise Rank-aware Interpolation Weights for Composed Visual Data Retrieval

This work revisits the efficacy of simple linear interpolation within an embedding space, and introduces SRAIN, the first framework that dynamically predicts query-specific interpolation weights, and achieves the best in composed video retrieval and matches the current state of the art in composed image retrieval.

Boseung Jeong, T. Park, Donghyeon Kwon et al. · 1 citation
Preprint Sep 2026

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constr...

Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar et al. · 0 citations
Preprint Sep 2026

Nearest but Not Dearest: Shared Curator-Feedback Infrastructure for Content-Only Search and Recommendation

A deployed B2B music-discovery platform serves both query-driven search (text prompts, vibe tags) and seed-driven recommendation (seed-track and artist stations) over one licensed catalog, one LAION-CLAP joint audio-text embedding space, one candidate-generation filter, and one ranking head -- and neither path consumes...

Matt Sandler · 0 citations
#artificial intelligence Preprint Sep 2026

Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval

Retrieval-augmented generation (RAG) depends on dense retrieval: each document is stored as a learned vector, and a query is answered by finding its nearest neighbors in that vector space. Keeping one full-precision vector per document is the dominant index cost at corpus scale, so retrieval systems replace each vector...

Pei-Chun Hua, Yun-Ming Xiao · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.