This work presents a tool that implements data-driven approaches that can dynamically select a smaller number of instances that provide sufficient statistical evidence to evaluate the relative performance of a given set of solvers and makes them readily accessible to solver developers, thus enabling them to obtain swifter feedback on their ideas.
Large language models (LLMs) have recently improved their problem-solving abilities and can solve complex mathematical problems with an increasing accuracy, necessitating the development of more challenging benchmarks. Over the years, the performance of LLMs on several benchmark datasets has also improved, motivating t...
Anurag Dutta, S. Priya, A. Ramamoorthy et al.· AppliedMath· 0 citations
AlgoWorlds is introduced, a benchmark that transforms formally specified combinatorial optimization problems into partially observed decision environments with verifiable global optima, and seven leading LLMs are evaluated, including Claude Opus 4.8 and GPT-5.6 Sol.
Despite decades of intensive research and optimization, modern Boolean Satisfiability (SAT) solvers have reached a plateau where significant performance gains are increasingly difficult to achieve. While Large Language Models (LLMs) have demonstrated remarkable capabilities in pattern recognition and code generation fo...
Mao Luo, Hang Ding, Chumin Li et al.· Proceedings of the Thirty-Fi...· 0 citations
Python is widely used in scientific research because it enables rapid development and provides rich ecosystems for data analysis, artificial intelligence (AI), and machine learning. However, customized research code can become prohibitively slow as experiments scale. This challenge is particularly acute in discrete-eve...
Combinatorial problems appear in numerous industrial applications. A common approach is to formulate these problems as declarative constraint models that can subsequently be compiled to and solved by a range of back-end solvers. Recent work shows that Large Language Models (LLMs) can produce correct models from natural...
: While Sudoku is governed by simple rules, it represents a complex constraint-satisfaction problem ideal for simulating and evaluating the performance of intelligent agents. This study presents a simulation framework designed to evaluate the behavioral effectiveness of various automated solvers. Deterministic methods,...
Vincent Sieso, Ludivine Lasserre, A. Doncescu· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.