Skip to content
Conference

LLM-Based Automated State Machine Generation and Assessment

Aug 2026 · 2026 IEEE 34th International Requirements Engineering Conference Workshops (REW) · pp. 291-300 · 1 citation · 36 references

Abstract

UML state machine modeling is a critical activity for the validation/verification of requirements and the specification of dynamic system behavior. Traditionally, state machines are manually crafted and assessed by experienced engineers based on natural-language requirements - a time-consuming and errorprone procedure. While many structured natural-language-based automated generation approaches exist, automated assessments have received little attention. In this paper, we empirically investigate the capabilities of state-of-the-art Large Language Models (LLMs) to (i) fully automate UML state machine generation from non-structured natural-language requirements, and (ii) automatically assess generated outputs by comparing them to a ground-truth solution. We first use human assessment to evaluate the generation quality of Claude Sonnet 4.5 using the singlestage baseline and a two-stage prompting approach. We then compare three LLMs - Claude Sonnet 4.5, GPT-5.5, and Gemini 3.1 Pro Preview - by evaluating their assessment performance (i) against human assessment and (ii) relative to each other for state machines without human assessment. Our experiments indicate that the two-stage prompting approach improves generation quality $(F_{1}$-score 0.898) over the single-stage baseline $(F_{1}$-score 0.775), driven by recall gains while precision remains very high across both configurations (> 0.91), confirming that the generation gap is one of coverage rather than correctness. Furthermore, our assessment comparison shows that all LLMs do not reach assessment levels required for full automation, exhibiting a systematic conservative bias by missing credit that human raters award rather than fabricating it. Based on these results, we suggest a hybrid assessment strategy. Our evaluation highlights both the potential and the limitations of current LLMs for automated state machine generation and assessment, providing a baseline for future research in this domain.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.