The underlying causes of LLM failures in OCL generation are investigated, framing the task as a graph reasoning problem over UML class diagrams and finding that OCL generation performance significantly degrades with increasing navigation depth and structural complexity.
Abstract
Large Language Models (LLMs) are increasingly used to generate Object Constraint Language (OCL) constraints from natural language specifications and UML class diagrams. However, existing work mainly focuses on improving accuracy, with limited understanding of why these models fail. Aims. This study investigates the underlying causes of LLM failures in OCL generation, framing the task as a graph reasoning problem over UML class diagrams. Method. We conduct an empirical evaluation using the PathOCL dataset across six state-of-the-art LLMs. We analyze the impact of UML structural properties (e.g., navigation depth and model complexity), lexical similarity, prompt ordering strategies, and graph-aware prompting on OCL correctness. Results. We find that OCL generation performance significantly degrades with increasing navigation depth and structural complexity. Lexical similarity has limited influence, while textual ordering of UML elements affects performance. Graph-based prompting yields partial improvements but does not eliminate structural reasoning errors. Conclusions. OCL generation is primarily constrained by graph reasoning limitations rather than purely linguistic factors. These results highlight structural reasoning as a key bottleneck for current LLMs in model-driven engineering tasks.
ClosureBench is introduced, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth with programmatically verified ground truth: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-verified correctne...
Large Language Models (LLMs) have shown promising performance in generating Object Constraint Language (OCL) constraints from natural language specifications. However, existing evaluations rely on publicly available UML models, which may overestimate generalization due to potential data leakage and reliance on recurrin...
Hamza Attarwala, Moataz Chouchen, Omar Alam et al.· Proceedings of the ACM/IEEE...· 0 citations
Evaluation of tool-augmented Large Language Models (LLMs) has not advanced far beyond final answer accuracy, and neglects in-depth evaluation of reasoning ability despite it being a central claim of recent models. We aim to address this gap by developing a dependency graph-based evaluation to give insight into meta-lev...
Nick Ferguson, Alan Bundy, Kwabena Nuamah· Proceedings of the Thirty-Fi...· 0 citations
Large language models (LLMs) have significantly advanced automatic code generation. Although they have demonstrated remarkable ability in writing standalone code, they still struggle with real-world repository-level coding tasks due to the lack of repository-specific knowledge. Existing studies typically employ retriev...
Tanghaoran Zhang, Xin-Jun Mao, Yu-Xin Zhao et al.· ACM Transactions on Software...· 0 citations
Software testing is a crucial process in software development. Despite its importance, the software testing process is often time-consuming and inconsistent when performed manually. With the rapid development of Large Language Models (LLMs), software testing processes such as generating test cases have utilized LLMs. N...
Shania Priccilia, Haryono Soeparno, F. L. Gaol et al.· International Conferences on...· 0 citations
Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, ra...
Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.