Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual rep...
Lucas Bandarkar, C. Peng, A. H. Ahmed et al.· 0 citations
This work presents a first study of how hybrid attention impacts the multilinguality of LLMs, showing that cross-lingual representations in hybrid models develop in patterns tied to the ordering of recurrent and full-attention layers.
Lucas Bandarkar, Jun-Ling Hu, Chen-Yuan Yang et al.· 0 citations
Six submissions under the team name ESTS to the unconstrained WMT26 Model Compression Shared Task for English--Simplified Chinese and English--Egyptian Arabic are described, describing how they use task-specific routing mass to rank experts and cross-lingual routing divergence to allocate retained capacity across layer...
Liu O. Martin, Lucas Bandarkar, Nanyun Peng· 0 citations
OmnilingualGAIA2 is introduced, a machine-translated expansion of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier, and it is argued that multilingual agentic evaluation must become a standard part of the reporti...
Andrea Caciolai, P. L. Cabot, Chierh Cheng et al.· 2 citations
This paper presents a method for aggressively pruning experts from modern mixture-of-experts LLMs while incurring negligible degradation in translation quality, and shows that translation requires only a fraction of the LLM, enabling substantial compression of the MoE blocks that contain over 90% of parameters.
Liu O. Martin, Lucas Bandarkar, Nanyun Peng· arXiv.org· 2 citations· ⚡1
There is potential to improve cross-lingual parametric knowledge transfer during post-training by providing the LLMs with the key entities of the questions in their source language and finding that this disproportionately improves cross-script questions.
Lucas Bandarkar, Alan Ansell, Trevor Cohn· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.