Skip to content

From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis

Jul 2026 · arXiv.org · Vol abs/2607.28307 · 0 citations · 44 references
Computer Science

TL;DR

This study investigates whether an LLM can bridge requirements engineering and architectural design, generating architectures solely from textual requirements and evaluating structural agreement and perceived quality of results, and shows potential for requirements-driven synthesis when guided by exemplar prompting.

Abstract

Microservice architectures have become dominant for modernizing monolithic systems, yet identifying appropriate services remains challenging and largely manual. Existing decomposition approaches are predominantly code-centric, limiting applicability in early design stages where only textual requirements are available. Despite advances in Large Language Models (LLMs), limited empirical evidence exists on their ability to synthesize complete microservice architectures from natural-language requirements, including service definitions and inter-service interactions. This study investigates whether an LLM can bridge requirements engineering and architectural design, generating architectures solely from textual requirements and evaluating structural agreement and perceived quality of results. We conduct a mixed-method study using OpenAI o3 under zero-shot (ZS) and few-shot (FS) prompting across two systems (Bookstore, PetClinic), one execution per system/condition. Architectures are evaluated through (i) comparison with reference architectures using precision, recall, and F1-score for service identification and communication recovery, and (ii) a blinded expert assessment of correctness, completeness, modularity, and plausibility, plus open feedback synthesis. OpenAI o3 identifies services with higher agreement under FS prompting (F1 = 0.79 for ZS versus = 0.97 for FS). Communication recovery is more challenging: ZS produces dense architectures with high recall but low precision (F1 = 0.61), while FS improves agreement, reaching F1 = 0.82 and reducing unsupported dependencies. Expert evaluation corroborates these results, with FS architectures perceived as more modular, coherent, and plausible than ZS outputs. OpenAI o3 shows potential for requirements-driven synthesis when guided by exemplar prompting. Results are model- and context-specific from two small systems, not model-independent proof.

View source

Similar papers

Jul 2026

Structural Validation of LLM-Generated Microservice Decompositions Using Source-Code Dependencies

The findings demonstrate that structural evaluations of LLM-generated decompositions should explicitly control for mapping coverage, as apparent differences between prompting strategies may otherwise reflect methodological bias rather than genuine architectural quality.

Daniel Silva, Renan Alves, Emanuel Dantas et al. · 0 citations
Preprint Aug 2026

Doc2CI: A Multi-Service Study of CI Configuration Generation Using Large Language Models

A large empirical study on using LLMs to generate CI configurations from natural language across services and model families suggests that similarity and validity are distinct objectives for CI generation and motivate schema-aware evaluation and tooling for LLM-based configuration generation.

T. A. Ghaleb · 0 citations
#artificial intelligence Review Sep 2026

Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study

A preliminary study using four LLMs with zeroshot and fewshot prompting on 30 developer-written ADDs from open-source projects finds that the generated ADDs are often too long, implementation-focused, and miss the rationale behind the decision.

Amey Karan, Rudra Dhar, Mohamed Soliman et al. · 0 citations
Conference Open access Sep 2026

Augmenting Enterprise Architecture With Large Language Models: A Single-Case Empirical Evaluation Of The TOGAF Preliminary Phase In PropTech

This research affirms the necessity of a human-in-the-loop paradigm in the EA discipline within this specific context and suggests that LLMs function effectively as augmentation accelerators rather than human replacements, thereby shifting the enterprise architect's focus from manual drafting to analytical validation a...

R. Darmawan, Alfa Yohannis · 0 citations
Book Open access Oct 2026

Bridging Feature Models and Users: LLM-Generated Low-Code Configurators for Software Product Lines

This work draws on principles from Low-Code Development (LCD) and introduces an LLM-driven approach for automatically generating interactive Web-based configurators directly from feature models, to simplify the development of configurators and support feature selection itself through low-code techniques that make varia...

Patrick Wegerer, Bernhard Schenkenfelder, Rudolf Ramler et al. · 1 citation
Review Open access Aug 2026

Large Language Models for Software Architecture Design Support in Self-Adaptive Systems: Early Insights from an Exploratory Systematic Review

An initial characterization of LLM-supported design in self-adaptive systems in SASS is contributed, research directions are outlined, and discussion within the community on advancing LLM-supported architectural design for self-adaptive and autonomous software systems is stimulated.

Nadeem Abbas, Nazia Shahzadi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.