Skip to content

Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

Jul 2026 · arXiv.org · Vol abs/2607.28979 · 4 citations · 41 references
Computer Science

TL;DR

Mixture-of-Translators (MoT), a cache translation framework that maps context KV caches from a source LLM into the cache space of a target LLM, is proposed, demonstrating scalable KV cache reuse across heterogeneous LLMs.

Abstract

Heterogeneous Large Language Model (LLM) systems increasingly rely on shared contexts, retrieved evidence, and multi-agent dialogue histories, yet their internal key-value (KV) caches remain model-specific and cannot be reused across architectures. Consequently, each model must repeatedly prefill or store caches for the same context, limiting the scalability of multi-model reasoning and long-context generation. We propose Mixture-of-Translators(MoT), a cache translation framework that maps context KV caches from a source LLM into the cache space of a target LLM. Unlike prior approaches that depend on a single projection path or global shared latent space, MoT uses multiple translator modules to capture diverse source--target mappings. To further reduce residual translation error, we introduce a Context Correction Loss that aligns the replayed target trajectory with the native target trajectory. We reveal two competing failure modes in cache translation: propagated translation shift from early injection and last-state shift from late injection. MoT addresses them through translator mixtures and target-side correction. Across homogeneous and heterogeneous translations among Qwen2.5, GPT-2, and OPT models, MoT preserves downstream QA performance, including Qwen2.5-7B-scale translation with 51.0% average closed-set QA accuracy and 0.43 average extractive QA F1. In practical case studies, MoT enables quality-preserving memory reuse for multi-agent reasoning and retains 96.3% of direct-context quality in long-context cache-augmented generation, demonstrating scalable KV cache reuse across heterogeneous LLMs.

View source

Similar papers

#machine learning Preprint Sep 2026

KV-Lingo: Learning KV-Cache Translators with Distillation

Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context: although the context has already been processed by one mode...

Valérie Castin, Keitaro Sakamoto, A. Filippova et al. · 0 citations
Preprint Aug 2026

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

PAMT is proposed, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning that improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.

Yongshi Ye, Biao Fu, Chongxuan Huang et al. · 0 citations
Preprint Aug 2026

Reasoning about In-Context Samples for Machine-Translation

A novel fragment-based reasoning framework is introduced in which the model first extracts parallel source-target fragments from retrieved similar exemplars, and uses these fragments as intermediate reasoning traces to produce the final translation.

Maxime Bouthors, J. Crego, François Yvon · 0 citations
#artificial intelligence Preprint Sep 2026

Towards Evolving Context Parameterization for Large Language Models

Context parameterization enables large language models (LLMs) to internalize contexts into reusable model parameters, avoiding repeated processing across subsequent queries. However, existing methods typically assume static contexts and lack explicit mechanisms for distinguishing validity states under continual updates...

Xiao Shi, Zhe-Rui Li, Yi-Ming Jiang et al. · 0 citations
Open access 2026

Large-Language-Model-Enhanced Adaptive Optimization for Data-Limited Multi-Domain Machine Translation

DKA-LLM-MT is proposed, a large-language-model-enhanced adaptive optimization framework for data-limited multi-domain machine translation that provides an effective and reliable solution for domain-sensitive machine translation under limited bilingual supervision and offers practical support for bilingual reading, spec...

Wei Yan · 0 citations
#artificial intelligence Preprint Sep 2026

Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may di...

Hasan Alkhder, Mohammad Abboush, I. Tchappi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.