Skip to content
Preprint

Why Do LLMs Fail at OCL Generation? A Graph Reasoning Perspective

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

The underlying causes of LLM failures in OCL generation are investigated, framing the task as a graph reasoning problem over UML class diagrams and finding that OCL generation performance significantly degrades with increasing navigation depth and structural complexity.

Abstract

Large Language Models (LLMs) are increasingly used to generate Object Constraint Language (OCL) constraints from natural language specifications and UML class diagrams. However, existing work mainly focuses on improving accuracy, with limited understanding of why these models fail. Aims. This study investigates the underlying causes of LLM failures in OCL generation, framing the task as a graph reasoning problem over UML class diagrams. Method. We conduct an empirical evaluation using the PathOCL dataset across six state-of-the-art LLMs. We analyze the impact of UML structural properties (e.g., navigation depth and model complexity), lexical similarity, prompt ordering strategies, and graph-aware prompting on OCL correctness. Results. We find that OCL generation performance significantly degrades with increasing navigation depth and structural complexity. Lexical similarity has limited influence, while textual ordering of UML elements affects performance. Graph-based prompting yields partial improvements but does not eliminate structural reasoning errors. Conclusions. OCL generation is primarily constrained by graph reasoning limitations rather than purely linguistic factors. These results highlight structural reasoning as a key bottleneck for current LLMs in model-driven engineering tasks.

View source

Similar papers

#machine learning Preprint Aug 2026

ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning

ClosureBench is introduced, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth with programmatically verified ground truth: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-verified correctne...

S. Goria · 0 citations
Book Open access Oct 2026

A Transformation-Based Benchmark for Evaluating the Robustness of LLMs in Generating OCL

Large Language Models (LLMs) have shown promising performance in generating Object Constraint Language (OCL) constraints from natural language specifications. However, existing evaluations rely on publicly available UML models, which may overestimate generalization due to potential data leakage and reliance on recurrin...

Hamza Attarwala, Moataz Chouchen, Omar Alam et al. · 0 citations
Conference Open access Sep 2026

Tool Call Dependency Graphs Enable Deep LLM Reasoning Evaluation and Better Explanations

Evaluation of tool-augmented Large Language Models (LLMs) has not advanced far beyond final answer accuracy, and neglects in-depth evaluation of reasoning ability despite it being a central claim of recent models. We aim to address this gap by developing a dependency graph-based evaluation to give insight into meta-lev...

Nick Ferguson, Alan Bundy, Kwabena Nuamah · 0 citations
Open access Oct 2026

LLM See, LLM Do: Enhancing Repository-Aware Code Generation in LLMs Through Logical Context and Usage Knowledge

Large language models (LLMs) have significantly advanced automatic code generation. Although they have demonstrated remarkable ability in writing standalone code, they still struggle with real-world repository-level coding tasks due to the lack of repository-specific knowledge. Existing studies typically employ retriev...

Tanghaoran Zhang, Xin-Jun Mao, Yu-Xin Zhao et al. · 0 citations
Conference Aug 2026

Semantic Tree: Leveraging Semantic UI Representation for Test Case Generation

Software testing is a crucial process in software development. Despite its importance, the software testing process is often time-consuming and inconsistent when performed manually. With the rapid development of Large Language Models (LLMs), software testing processes such as generating test cases have utilized LLMs. N...

Shania Priccilia, Haryono Soeparno, F. L. Gaol et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs

Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, ra...

Ramya Keerthy Thatikonda, W. Buntine, Ehsan Shareghi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.