Skip to content

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

Sep 2026 · 0 citations · 27 references
Computer Science Engineering

TL;DR

The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, enabling the LLM to behave as an audio language model (ALM) and improves the scalability of ALMs.

Abstract

This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.

View source

Similar papers

#natural language process... Preprint Sep 2026

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M parameters. We introdu...

Kai-Yang Li, Shaobo Han, Yue Tian et al. · 0 citations
#natural language process... Preprint Sep 2026

SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a...

H. Le, L. Nguyen, Minh Tri Dao · 1 citation
#artificial intelligence Preprint Sep 2026

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short beha...

Hao-Jun Zhang, Yi Zou, Min Chen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

This work presents TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks and introduces the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the eval...

Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh et al. · 2 citations
Conference Open access 2026

DeepSeek-V3: Architecture and Optimizations-A Practical Review

The design of transformer-based Large Language Models (LLMs) is being radically changed through new architectures that are able to overcome scalability limitations of previous designs, including Mixture-of-Experts (MoE), Multi-Head Latent Attention (MLA), and Multi-Token Prediction (MTP). As an open-weighted model rele...

Yassine Zouhdi, B. Hdioud · 0 citations
Preprint Aug 2026

Token Communication for Multimodal Large Language Model

With the broad success of the Transformer architecture, token is becoming a new basic information processing unit. This trend is especially evident in multimodal large language models (MLLMs), where both visual and textual information are represented and processed as tokens. With the rapid deployment of MLLMs, the effi...

Jing-Kai Ying, Zhi-Jin Qin, Yuan Shen et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.