Oct 2026· Proceedings of the ACM/IEEE 29th International Conference on Model Driven Engineering Languages and Systems· 0 citations· 12 references
TL;DR
This paper proposes an approach that reformulates the generation problem as a code generation task, on which the LLMs excel, and evaluates the approach along four dimensions—scalability, consistency, diversity, and realism—across two use cases and three LLMs.
Abstract
The automated generation of instance models plays a central role in object-oriented software testing, benchmarking of graph databases, and the assurance of cyber-physical systems. Model generators are tools designed to address this problem, typically taking as input a set of constraints that the generated models must satisfy. Traditional model generators rely on SAT solvers, search-based techniques, grammar-based approaches or learning-based generative techniques. In this paper, we investigate the potential of Large Language Models (LLMs) as a new paradigm for instance model generation. We propose an approach that reformulates the generation problem as a code generation task, on which the LLMs excel. Specifically, meta-models are encoded as Pydantic models, while well-formedness constraints are expressed using data validation Pydantic constraints. The LLM is prompted to generate executable code that constructs valid instance models, and a feedback loop is employed to iteratively correct invalid outputs. We evaluate our approach along four dimensions—scalability, consistency, diversity, and realism—across two use cases and three LLMs, and compare it against Refinery, a state-of-the-art model generator. The LLM-based approach demonstrates strong scalability with respect to instance model size, being able to generate models exceeding 2000 elements by effectively leveraging programmatic constructs such as loops. In terms of consistency, one of the evaluated LLMs achieves a high probability of generating fully consistent models, even for very large instances. However, diversity emerges as a major limitation of LLM-based generation, with the proposed generator showing a significant drop in diversity as the scope increases. Finally, while the LLM-based generator exhibits a certain degree of realism, its performance in this dimension is influenced by the domain of the meta-model.
BRIDGE is presented, a structured prompting framework that decomposes verification into three interconnected domains: Code (implementations), Specifications (formal intent), and Theorem State-ments (constructive correctness claims), and elicits domain-specific intermediate reasoning to connect them.
Robert Joseph George, Carson Eisenach, Udaya Ghai et al.· 0 citations
Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...
Large Language Models (LLMs) have shown promising performance in generating Object Constraint Language (OCL) constraints from natural language specifications. However, existing evaluations rely on publicly available UML models, which may overestimate generalization due to potential data leakage and reliance on recurrin...
Hamza Attarwala, Moataz Chouchen, Omar Alam et al.· Proceedings of the ACM/IEEE...· 0 citations
This work introduces Path2Spec, a divide-and-conquer framework that leverages LLMs to extract all execution paths from an input program, generates path-specific specifications for each, and merges them into a comprehensive overall specification.
Dan Huang, Zhensu Sun, Hui-Hui Huang et al.· 0 citations
Automated unit test generation promises to reduce the cost of software quality assurance, and hence, is attracting attention from both academia and industry. Yet, generating assertions that are executable, meaningful to developers, and able to catch faults remains an unsolved challenge. Existing approaches either rando...
Jia-Lun Cao, Hao-Yu Wang, Hao-Ran Yan et al.· Proceedings of the ACM on So...· 0 citations
This paper extends the state-of-the-art FlowRepair approach by replacing a subset of its mutation operators with LLM-generated mutations, enabling more flexible and expressive patch generation and highlighting fundamental limitations of naively integrating LLMs into search-based APR.
Ayesha Irshad, Pablo Valle, J. Ayerdi et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.