PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Hongliang Li, Yijin Liu, Zhiwei Zhang et al.
· 1 citation
2 papers indexed here
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
Benchmark evaluations reveal that agent performance varies substantially across languages and drops sharply on the harder cross-lingual tasks, and analysis shows that multilingual execution exposes systematic failure modes across planning, tool interaction, and decision-making in long-horizon agents.