#natural language process...
Mar 2026
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
This work proposes Hikari, a policy-free, end-to-end model for simultaneous speech-to-text translation and streaming transcription, and presents a supervised fine-tuning strategy that trains the model to recover from delays, significantly improving the quality-latency trade-off.
Roman Koshkin, Haesung Jeon, Lian-Bo Liu et al.
· arXiv.org · 2 citations