2026· Proceedings of the 21st International Conference on Software Technologies· 0 citations· 24 references
TL;DR
The proposed model aims to support the formalization of model selection processes, improve decision-making, and enhance the traceability and transparency of LLMOps practices and forms part of a broader research effort toward the formalization of the entire LLMOps life cycle.
Abstract
: The efficient adoption of Large Language Models (LLMs) in enterprises requires formalizing and systematically implementing LLM Operations (LLMOps), as it could reduce costs and ensure the reproducibility of LLM solutions. One of the critical phases of LLMOps is selecting a suitable LLM, as it can significantly affect a system’s performance, feasibility, and maintainability during the operational phase. As a result, uninformed decisions about model selection that do not align with business, functional, and non-functional requirements may severely impact subsequent stages of the LLMOps life cycle. To investigate the model selection phase, we conducted an extensive Systematic Literature Review (SLR) and synthesized its findings to derive a set of activities fundamental to this phase. These activities are presented in an end-to-end UML activity diagram. The proposed model aims to support the formalization of model selection processes, improve decision-making, and enhance the traceability and transparency of LLMOps practices. This work forms part of a broader research effort toward the formalization of the entire LLMOps life cycle, with the goal of facilitating scalable and reliable integration of LLMs in enterprise environments.
: Large Language Models (LLMs) are increasingly used to evaluate software engineering artifacts. This paper investigates the reliability of LLMs in evaluating UML diagrams generated through reverse engineering processes (source code). We ask: do LLM assessments align with those of human experts? A COMAS-based framework is proposed for trust-aware LLM selection. In the original COMAS, indirect trust is computed via Dijkstra max-product chains. We replace this with direct human–LLM comparisons for each criterion independently. This is valid since LLMs can evaluate any diagram on demand. Trust is measured using an L1-norm-based divergence that is robust under small-sample conditions. The developed framework supports flexible per-criterion LLM selection via expert-defined thresholds and weights. An empirical study involved 30 students evaluating five UML sequence diagrams. Four evaluation criteria were used: completeness of requirements description, non-contradiction to requirement specification, UML notation compliance, and Miller’s principle. Four LLMs were compared: Groq, Gemini, Mistral, and Copilot. For 90% of participants, a combination of two or three LLMs outperformed any single model. Hybrid per-criterion LLM selection is a viable strategy for automated UML diagram evaluation.
Olena Chebanyuk, Carles Sierra· Proceedings of the 21st Inte...· 0 citations
The use of large language models (LLMs) to support non-modeling experts of multi-perspective EM are investigated and LLMs can be seen as assistive technology for certain tasks in EM.
Peter-Alexander Kolev, Hauke Hansen Pruss, J. Wilken et al.· Journal of Software and Syst...· 0 citations
Large Language Models (LLMs) are emerging as a promising tool in Business Process Management for comparing and validating process models. In this study, we evaluated an LLM’s ability to compare reference models with systematically modified variants representing typical modeling mistakes as well as harmless variations, such as layout or wording changes. The results show that the applied LLM can reliably detect structural and semantic differences between formal business process models using Business Process Model and Notation, while distinguishing them from acceptable variations, demonstrating strong potential for automated model validation. However, the LLM’s performance and accuracy are influenced by factors such as model complexity, the number of inserted modifications, and the total number of models and modifications provided simultaneously. High reliability is achieved when models are presented in a standardized, semi-structured format and supported by clear prompting instructions. Even multiple models can be processed effectively, up to a certain threshold of total modifications. Overall, the findings suggest that generative Artificial Intelligence tools for natural language processing, such as LLMs, may provide meaningful support in process model validation, offering efficiency gains and a level of abstraction that exceeds manual comparison.
Christian Bennoit, S. Zamani, Tobias Greff· Process Science· 0 citations
Business process management (BPM) faces three persistent structural challenges: (1) the exclusion of non-technical stakeholders from process design due to the complexity of BPMN 2.0 notation; (2) the fragmentation among requirements engineering, user interface design, and process modeling; and (3) the absence of a transparent mechanism for managing stakeholder intellectual property (IP) rights. This paper introduces EffABPMN2MV — a ten-layer collaborative design ecosystem operating within an immersive metaverse environment. The framework comprises three integrated components: (i) EffABPMN2MV-ML, a modeling language that maps metaverse UI interactions to BPMN 2.0 elements; (ii) EffABPMN2MV-Tools, a GPT-4-powered AI engine that automatically generates four standardized software deliverables; and (iii) EffABPMN2MV-SCB, a Hyperledger Fabric smart contract subsystem that immutably records stakeholder contributions and allocates LNT participation tokens. A controlled experiment with two parallel nine-member teams (Group A: EffABPMN2MV; Group B: Scrum) and an independent expert panel (n=15) was conducted using a banking loan application case study. EffABPMN2MV reduced documentation time by 58.3% (t=4.82, p<.001, Cohen's d=2.28), improved document quality by 21.4% (d=1.65), enhanced stakeholder experience (d=2.27), and raised legal transparency from 2.58 to 4.61 out of 5.0 (d=3.47). All effects were statistically significant (p<.001) with large effect sizes (d>1.0). EffABPMN2MV represents a theoretically grounded and empirically validated contribution to BPM, collaborative software design, and blockchain-enabled information systems.
Masoud Rezaei, Abbas Mirzaei, Babak Nouri-Moghaddam et al.· Journal of Resource Manageme...· 0 citations
Interviews with sixteen early-adopter software professionals who integrated LLM-based tools into their day-to-day work in early to mid-2023 offer actionable implications for developers, organizations, educators, and tool designers seeking to integrate LLMs responsibly into professional software practice.
Benyamin T. Tabarsi, Heidi Reichert, Sam Gilson et al.· Empirical Software Engineeri...· 21 citations· ⚡1
The automated generation of business process models from natural language descriptions has recently attracted growing attention at the intersection of Business Process Management (BPM), Natural Language Processing (NLP), and Large Language Models (LLMs). This paper presents a Systematic Literature Review (SLR) on the current state of research in this emerging field. Following the guidelines of Kitchenham et al. and the PRISMA framework, 29 studies published between January 1st, 2023 and March 10th, 2026 were identified, selected, and analyzed. The review addresses two research questions focusing on the applied methodological approaches, the used LLMs, the employed modeling languages, as well as the evaluation strategies, challenges, and limitations reported in the literature. The results show a clear shift from traditional NLP-based techniques toward LLM-only and hybrid approaches. OpenAI’s GPT family, especially GPT-4 and its variants, dominates the field, while BPMN is by far the most frequently used target process modeling language. Furthermore, existing studies evaluate automated process model generation primarily through output-focused methods, such as quantitative metrics, expert reviews, and comparisons with alternative or human-created process models. At the same time, the reviewed studies reveal important challenges, including the continued need for human involvement and the output quality. Overall, current approaches show strong potential, but they still act more as intelligent assistants than as fully autonomous process modelers.
L. F. Hörner, Maximilian Möller, Manfred Reichert· IEEE Access· 0 citations