Skip to content

Last Translation Benchmark

Vilém Zouhar Niyati Bafna Mukund Choudhary Maike Züfle Sara Rajaee Pinzhen Chen Jannis Vamvas Sara Papi Ona de Gibert Bhavitvya Malik Eliya Habba Orfeas Menis-Mastromichalakis Patrícia Schmidtová Michelle Wastl Sheriff M. Issaka Leshem Choshen Stella Biderman Antonis Anastasopoulos Jan Niehues Rico Sennrich Mrinmaya Sachan Ondrej Bojar Kenton Murray Jörg Tiedemann Alham Fikri Aji Philipp Koehn C. Monz Alexandra Birch Sowmya Vajjala C. Kranti Cristina España-Bonet Nobin Sarwar David Kaczér Sourajit Saha Jonathan Tonglet Shunta Asano Malik Marmonier Daban Q. Jaff Vaisakhi Mishra H. Al-Khalifa Gabriele Sarti N. Rehlinger J. C. Villa Dominik Machácek Saugata Purkayastha Jagannathan Ramanujam Shubhashis Roy Dipta A. Nigam S. S. Yusuf Heejin Do J. Yahav Johannes-Rudolf David M. Staiano Zuzana Nadova Fred Philippy Ron Keinan Maria Lymperaiou Silvia Casola Fabian Retkowski Andrés Jerez Hanna Yukhymenko Sangwon Ryu Avantica Vempati Sukannya Purkayastha Adrian Cosma Erivan Inan V. Babenko W. Aissa Valentin Scourneau Fatima Haouari Venkata Gummadi Mehdi Jafarzadeh Manon Reusens Kaiser Sun Lukas Edman Shaomu Tan Giuseppe Gallipoli Pawan Sasanka Ammanamanchi Manar Ali Mohammad Sadegh Gholizadeh Dipankar Srirag Marek Suppa Javier García Gilabert R. Binkytė Ana-Maria Bucur Sabry E. Farrag Youssef Saber Yihong Liu Theresia Veronika Rampisela Christian Hoang N. Cojocaru Jan Koco'n Jean Maillard Xiao-Chuang Yuan Sina Ahmadi Daryna Dementieva Philipp Mondorf Kaustubh D. Dhole Valmik Nahata Roman Wixinger Amir Hossein Yari Shen-Bin Qian Manuel Tuor F. M. Thoker S. Troshin Lance Calvin Lim Gamboa Amir Arsalan Rezapour Kätriin Kukk Koel Dutta Chowdhury Shaswati Saha Seth Aycock Bo Chen Linh Vu Vatsal Venkatkrishna Shayan Bali Arafat Ahsan L. Nguyen Hassan Soliman Ngoc Quynh Tram Do Azmine Toushik Wasi R. Damanhuri M. Huber Kazuki Egashira J. Layacan D. Africa Vladislav Poritski Mike Zhang D. Shah Abdulaziz Nura Kani Luis Frentzen Salim Paul Gavrikov Bello Bello A. Munot A. Garg Y. Xavier Qiao-Yuan Zheng Kawsar Ahmed Debanshu Das Zi-Mu Wang G. Rao Kamile Dementaviciute Farhan Farsi L. D. M. S. Sai Teja Da-Wei Zhu Yi Fan Wei Liu M. Gaido Elias Herranen Sankalan Pal Chowdhury Karen Sanchez Guy Kaplan Farzad Shami Ashok Urlana Amir Hossein Kargaran Sofie Goethals Priyaranjan Pattnayak O. Volchek Marii Ojastu Hongbin Na Emilian Radoi Chen-Yi Zhao Carlos Hinojosa Andrei Niculae Andrea Gregor de Varda Zaid Alyafeai Tomasz Limisiewicz Reem Alzahrani Pouya Sadeghi Nehal Kathrotia Mateusz Lango Enzo Doyen Alex Flückiger Ulysses Sekai Tully Carr Samuel Simko R. Tiwari Rishit Dagli Isaac Caswell Bo-Wen Yi Aicha Chorana Selja Keränen Sadiksha Chitrakar Muhammad Ravi Shulthan Habibi Joy Olusanya Bishal Shrestha Zhengxiang Wang Vivek Harsha Lakkamaneni Sophia Conrad Panayiotis N. Panayiotou Nazia Tasnim Marta Punsola Munárriz Marko Čuljak Luis Lara Jenny Chim Jannatul Nayem Fidel Rodríguez Velásquez Eran Yahav Blanka Kövér Beatrice Savoldi Anmol Goel Aishik Mandal Tosin P. Adewumi Raoyuan Zhao Mykola Haltiuk Antonia Karamolegkou Yu-Xing Lu Thura Aung N. Al Mousa Tommaso Cerruti Raia Abu Ahmad Béni Egressy Alireza Pakniat Stéphane J. P. S. Thunus Rachel Bawden Lena Libon Samridhya Biswas Prakhar Gupta Nusrat Jahan Lia T. Nguyen Natchapon Jongwiriyanurak Minh N. Do I. Barač D. Kuzmin Badal Nyalang Antoine Taroni Andy Catruna Rushikesh Zawar Roland C. Aydin Pavel Stepachev Ilai Yaron Levy Andreas Simons Rayyan Merchant Zi-Yi Yang Samuel Frontull Kenneth C. Enevoldsen Harris Abdul Majid T. Graf Tatiana Bielakova Sharifa Djurabaeva Shao Ji Jirui Qi Ayla Rigouts Terryn Yurii Paniv Xi-Yang Fu Sunisth Kumar S. Satish Papa Abdou Karim Karou Diallo Mengyu Ye Maximilian Vieweg Matej Akrap Kristýna Onderková Joseph Attieh Ivan Yuri De Leon Ibrahim Baroud Esrael Teferi Tensay Elisabeth Fittschen David Dukić Benoît Sagot Jingwei Ni Yu Fan Juri Opitz
Sep 2026 · 0 citations · 87 references
Computer Science

Abstract

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, opaque, and vulnerable to reward-hacking. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18
#computer vision Open access Mar 2014

Happy software developers solve problems better: psychological measurements in empirical software engineering

A study with 42 participants investigates the relationship between the affective states, creativity, and analytical problem-solving skills of software developers and offers support for the claim that happy developers are indeed better problem solvers in terms of their analytical abilities.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 216 citations · ⚡13
#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.