Open access
Jul 2026
GMTW-Ro: a deterministic benchmark for evaluating large language models on grounded Romanian tasks
GMTW-Ro is introduced, a benchmark designed to evaluate whether large language models can reliably follow complex instructions in Romanian, rather than merely produce fluent text, and raises important questions about how current language adaptation pipelines preserve instruction-following and structured reasoning capabilities.
Andrei-Ștefan Bulzan, Bogdan Morariu, Andrei-Razvan Joldea et al.
· Frontiers in Artificial Inte... · 0 citations