Skip to content

Structured PREreview of "Automating Business Intelligence Requirements with Generative AI and Semantic Search"

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record is a permanently preserved version of a Structured PREreview. You can view the complete PREreview at https://prereview.org/reviews/22877617. Does the introduction explain the objective of the research presented in the preprint? Yes Yes. The introduction explicitly explains the objective of the research presented in the preprint. Specifically, the introduction details the following points: Problem Context: Traditional methods for capturing and managing evolving Business Intelligence (BI) requirements are labor-intensive, error-prone, and require extensive coordination among data analysts, subject matter experts, and business stakeholders. This leads to gaps between business needs and technical implementations, repeated design cycles, and increased risk during cloud or technology migrations. Core Research Objective: To solve these challenges, the authors propose AUTOBIR, a novel no-code system that combines Large Language Models (LLMs) and Semantic Search to automate and accelerate the specification of BI requirements. Scope and Deliverables: The introduction outlines that AUTOBIR uses a conversational interface to convert natural language inquiries into executable queries, analytics specifications, data dependencies, prototype code, and test-case reports. Beyond presenting the system, the research aims to detail its underlying architecture and discuss the broader potential of Generative AI for scaling enterprise data engineering processes. Are the methods well-suited for this research? Somewhat appropriate Justification: The methodology presented in the paper is well-suited for designing and building an automated Business Intelligence (BI) requirement elicitation system, following AI and software engineering best practices through most of the research. Methodological Strengths: Technical System Design: Combining Web Ontology Language (OWL) and R2RML bindings with vector database semantic search (Milvus/Pinecone) provides a rigorous approach to schema abstraction and sub-ontology retrieval . This directly addresses schema complexity before passing prompts to the Large Language Model (LLM). Robust Error Handling: The multi-layered self-debugging framework—incorporating syntax, semantic, and execution-based checkers—ensures generated queries are validated iteratively before deployment. Design Science Research (DSR): The authors employ DSR methodology to iteratively evaluate and refine the system based on real-world client engagements. Methodological Limitations: Qualitative Evaluation vs. Rigorous Benchmarking: The primary evaluation relies on qualitative feedback from structured sessions with 23 Subject Matter Experts (SMEs) across four domains, rather than a formal, controlled quantitative user study measuring task completion speeds or error rates against traditional methods. Generalizability and Scale: While benchmarked on standard datasets like AdventureWorks2014, Spider, and BIRD, the paper acknowledges threats to external validity regarding full scalability on massive real-world enterprise schemas (such as Spider 2.0 workflows). Overall, the combination of semantic modeling, LLMs, automated debugging, and expert key-informant feedback provides a solid, well-executed foundation for drawing valid conclusions about the tool's effectiveness, despite the need for a larger quantitative user study in future work. Are the conclusions supported by the data? Somewhat supported Justification: The conclusions drawn by the authors are reasonable and mostly supported by the evidence presented in the paper, though certain broad assertions rely on qualitative feedback rather than comprehensive quantitative measurement. Where the Data Supports the Conclusions: Functional Proof-of-Concept: The running example on the Microsoft AdventureWorks2014 benchmark dataset demonstrates that the system successfully processes natural language inquiries, maps them to sub-ontologies, and generates valid SQL queries, natural language explanations, execution outputs, and visualizations. Domain Adaptability: The system was deployed and tested across four distinct client domains (Security, Air Defense, Retail, and Banking), gathering qualitative feedback from 23 Subject Matter Experts (SMEs). This confirms that the modular architecture can adapt to varied metadata structures and business contexts. Benchmark Alignment: Comparative evaluation against standard datasets (Spider and BIRD) demonstrates that AUTOBIR's semantic pruning and self-debugging mechanisms effectively address common Text-to-SQL errors. Where Support is Limited: Lack of Quantitative User Performance Data: The paper concludes that AUTOBIR reduces labor costs and accelerates BI requirement delivery. However, these claims are supported primarily by qualitative key-informant feedback rather than controlled quantitative metrics (such as task completion time, error rates, or developer velocity comparisons against traditional methods). The authors explicitly note that a formal large-scale user study is planned for future work. Enterprise-Scale Validation: While the architecture is designed for distributed systems, the authors acknowledge that full scalability to massive enterprise schemas (such as those represented in Spider 2.0 with 1,000+ columns) remains to be empirically validated in large-scale production environments. In summary, the authors provide a realistic and grounded interpretation of their prototype and qualitative evaluation, avoiding extreme overreach while candidly acknowledging the empirical studies needed in future work. Are the data presentations, including visualizations, well-suited to represent the data? Somewhat appropriate and clear Justification: The paper presents a mixed set of visual materials. While the architectural and conceptual diagrams are clear and informative, the sample data visualization generated by the system and showcased in the paper suffers from significant usability and accessibility flaws. Strengths (Architectural & Schema Visualizations): System Architecture Diagram (Figure 3): Clearly delineates the distinction between offline Setup Tools (OntoDis, OntoManager, OntoSearch) and online Run-Time Tools, making the data flow easy to follow. Grounding View / Sub-Ontology Graph (Figure 1): Effectively visualizes knowledge graph nodes and object relationships, helping users understand how entities like Product, Currency, and SalesOrderHeader are linked. Weaknesses (Generated Data Visualizations): Inappropriate Chart Selection for High Cardinality (Figure 2): To represent the query results across 198 entities (total earnings per product), the paper showcases an automatically generated pie chart. Pie charts are fundamentally ill-suited for high-cardinality categorical data. Accessibility and Readability Barriers: In Figure 2 (right), dozens of narrow slices lead to severe visual clutter, unreadable overlapping text percentages, and an overcrowded, truncated legend. Lack of Formatting Best Practices: For 198 product items, standard data visualization best practices dictate using a sorted horizontal bar chart, a top-N filter, or an interactive table rather than a multi-slice pie chart. While the conceptual diagrams effectively communicate the system's engineering design, the primary example of an output data visualization in Figure 2 demonstrates clear readability and accessibility limitations. How clearly do the authors discuss, explain, and interpret their findings and potential next steps for the research? Very clearly Justification: The authors provide an exceptionally clear, structured, and insightful discussion of their qualitative findings, domain lessons, system limitations, and future research directions. Key Strengths of the Discussion and Next Steps: Comprehensive Lessons Learned (Section V): The paper offers a deep exploration of takeaways from deploying AUTOBIR across four distinct operational domains (Security, Air Defense, Retail, and Banking) with 23 Subject Matter Experts. It thoroughly analyzes critical topics such as the necessity of logical semantic layers over cryptic physical schemas, the role of human-in-the-loop interactive feedback, and key LLM security risks (including prompt injection, unauthorized data exposure, and query execution safety). Transparent Threats to Validity (Section VII): The authors dedicate an entire section to systematically evaluating internal, external, and construct validity threats. They candidly acknowledge current limitations—such as reliance on qualitative key-informant feedback rather than quantitative metrics, potential data selection biases, and schema scalability bounds. Explicit and Actionable Next Steps (Sections V, VII, VIII): Potential next steps are clearly defined rather than listed as vague generalizations. The authors specifically outline upcoming work, including: Conducting a formal, large-scale quantitative user study to measure developer velocity, task completion time, and labor cost reductions. Benchmarking the system against enterprise-scale real-world workflows like Spider 2.0 (with schemas exceeding 1,000 columns). Developing enhanced ontology alignment techniques for semi-structured and unstructured data sources. Overall, the authors demonstrate high clarity and critical depth in interpreting what their prototype achieves today and detailing the exact roadmap required for enterprise-grade deployment. Is the pr

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

Microsoft Research Blog Aug 12, 2026

MindTopo reveals VLMs’ spatial reasoning abilities

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research.

Microsoft Research Blog Jul 30, 2026

Echoverse: Deep, evolving environments for computer-use agents

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.