Skip to content
Open access

Models, Prompts, and Code! A Semi-Formal State Machine Language for Multi-Paradigmatic Software Development

2026 · International Conference on Software and Data Technologies · pp. 203-210 · 0 citations · 22 references
Computer Science

TL;DR

Compared to both classical UML tooling and fully LLM-based generation, the approach offers stronger determinism, better traceability, lower cognitive modeling effort, and reduced computational cost, while retaining the flexibility to express complex action behavior in natural language where formal specification would be unnecessarily burdensome.

Abstract

: Model-Driven Software Engineering has long excelled at generating code from static structural models, yet the specification and generation of dynamic behavioral models remains a persistent challenge. Meanwhile, Large Language Models (LLMs) offer flexible, natural-language based code generation but suffer from non-determinism and hallucinations. This paper presents a semi-formal approach that bridges these two paradigms for behavioral modeling via UML state machines. We contribute a textual modeling language that captures the essential elements of UML state diagrams—states, transitions, events, guards, and entry/exit actions—alongside a deterministic code generator that transforms state machine models into Java code following the Gang of Four State design pattern. The language supports two complementary action annotation styles: direct code fragments for concise, self-contained actions, and natural language descriptions for semantically richer behavior to be completed by an LLM weaver. LLM involvement is deliberately scoped to small, well-constrained action bodies, reducing token consumption and non-determinism compared to fully LLM-based approaches. Validated through the Gumball Machine case study, correctness is confirmed by automated tests covering state and transition coverage criteria, and repeating the LLM weaving step produced consistent results across all runs. Compared to both classical UML tooling and fully LLM-based generation, the approach offers stronger determinism, better traceability, lower cognitive modeling effort, and reduced computational cost, while retaining the flexibility to express complex action behavior in natural language where formal specification would be unnecessarily burdensome.

Read PDF

Similar papers

Review Sep 2026

Test-Driven Approaches to Software Engineering with Large Language Models: A Survey of Phases, Tasks, and Agent Skills

Tests increasingly participate in the decisions made by large language models and software engineering agents. They specify intended behavior, guide program construction and repair, select candidates, constrain transformations, and provide execution evidence for software analysis. These uses draw on test-driven develop...

Yun-Hao Liang, Cheng-Guang Gan, Rui-Xuan Ying et al. · 0 citations
Open access Sep 2026

Generative File Systems: Specification-Driven Synthesis and Evolution via LLMs

File systems are critical OS components that require constant evolution to support new hardware and emerging application needs. However, the traditional paradigm of developing features, fixing bugs, and maintaining the system incurs significant overhead, especially as systems grow in complexity. This paper proposes a n...

Qing-Yuan Liu, Heng Zhang, Mo Zou et al. · 0 citations
Jul 2026

Specula: Scaling formal specifications for autonomous model checking of system code

Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding. Specula employs large language model (LLM) based coding agents to autonomously develop TLA+ specifications, including...

Q. Cheng, Saad Mohammad Rafid Pial, Ruize Tang et al. · 0 citations
Preprint Sep 2026

Where the LLM Ends and Reliable Decisions Begin

Systems that turn natural-language descriptions of optimization problems into solver-ready code generally use a language model at every stage, including the final translation from a mathematical formulation into executable model-building code. We propose the ANVIL compiler architecture, where we separate these concerns...

Priyadarshan Patil, A. Basu, Vikas Reddy et al. · 0 citations
Review 2026

Using large language models to generate executable BPMN models based on text descriptions: an overview of approaches, limitations, and validation methods

It is argued that edge (sequence flow) generation is the weakest link once nodes are fixed, and typical structural failure modes (dangling nodes, disconnects, gateway violations, etc.) and causes tied to autoregressive generation are summarized.

Gennady G. Bulgakov, S. Yarushev · 0 citations

BRIDGE: Building Representations in Domain-Guided Program Synthesis

BRIDGE is presented, a structured prompting framework that decomposes verification into three interconnected domains: Code (implementations), Specifications (formal intent), and Theorem State-ments (constructive correctness claims), and elicits domain-specific intermediate reasoning to connect them.

Robert Joseph George, Carson Eisenach, Udaya Ghai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.