Skip to content

HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning

Jan 2026 · arXiv.org · Vol abs/2601.10187 · 0 citations · 45 references
Computer Science

TL;DR

This work proposes Homura, a reinforcement learning framework that explicitly optimizes the trade-off between semantic preservation and temporal compliance, and demonstrates that Homura significantly outperforms strong baselines, achieving precise length control that respects linguistic density hierarchies without compromising semantic adequacy.

Abstract

Large Language Models (LLMs) have achieved remarkable strides in multilingual translation but are hindered by a systemic cross-lingual verbosity bias, rendering them unsuitable for strict time-constrained tasks like subtitling and dubbing. Current prompt-engineering approaches struggle to resolve this conflict between semantic fidelity and rigid temporal feasibility. To bridge this gap, we first introduce Sand-Glass, a benchmark specifically designed to evaluate translation under syllable-level duration constraints. Furthermore, we propose Homura, a reinforcement learning framework that explicitly optimizes the trade-off between semantic preservation and temporal compliance. By employing a constrained reinforcement learning objective featuring a novel dynamic syllable-ratio reward, Homura effectively"tames"the output length. Experimental results demonstrate that Homura significantly outperforms strong baselines, achieving precise length control that respects linguistic density hierarchies without compromising semantic adequacy.

View source

Similar papers

Preprint Sep 2026

Beyond Saying Less: Fine-Grained Alignment for Informative and Faithful Vision-Language Models

This work proposes a fine-grained alignment framework that couples dense reward signals at the data level with precise credit assignment at the algorithmic level to prevent response-level shared advantages from allowing local hallucinations to compromise all other valid outputs within the same response.

Xing-Ming Long, Jie Zhang, Yue-Cong Min et al. · 0 citations
#natural language process... Preprint Sep 2026

DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation

Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning...

Ya-Yue Deng, Ding-Dong Wang, Yuxuan Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VAA-CSEC: Vote-guided Advantage Allocation for Chinese Semantic Error Correction

Chinese Semantic Error Correction (CSEC) targets semantic errors in Chinese text, which are typically more subtle and complex than spelling and grammatical errors but remain relatively underexplored. Existing LLM-based approaches face two recurring obstacles in this task: over-correction, and unclear interaction betwee...

Yi Han, Nan-Kai Lin, Juan Luo et al. · 0 citations
Conference Aug 2026

A Distributed Learning Approach for Controlled Grammar Transfer using Structure–Timbre Disentangled Audio Synthesis

This research article introduces a probabilistic framework for controlled grammar transfer in audio synthesis, tackling the issue of blending audio styles in a way that is both interpretable and disentangled. Current methods either depend on low-level feature matching or use neural models that mix structural and timbra...

Sridhar Varma K, Karthick S · 0 citations
#artificial intelligence Preprint Sep 2026

Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may di...

Hasan Alkhder, Mohammad Abboush, I. Tchappi et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.