Bench of Euler: A Benchmark for Evaluating the Problem-Solving Abilities of Large Language Models
Large language models (LLMs) have recently improved their problem-solving abilities and can solve complex mathematical problems with an increasing accuracy, necessitating the development of more challenging benchmarks. Over the years, the performance of LLMs on several benchmark datasets has also improved, motivating t...