Skip to content
Review

Scalable ETL Pipeline Architectures for Real-Time Transaction Analytics: Bridging Data Engineering and Business Operations

2025 · International Journal of Multidisciplinary Research and Growth Evaluation · 0 citations

TL;DR

The study concludes that real-time analytical performance must be evaluated through both engineering and operational outcomes, and recommends use-case-driven architectural selection, resilient hybrid deployment, embedded security and governance, automated quality assurance, transparent AI-assisted pipeline management and stronger collaboration between technical and business teams.

Abstract

This review examines the architectural, technological and organisational conditions required to develop scalable extract, transform and load pipelines for real-time transaction analytics. Its purpose is to clarify how contemporary data-engineering capabilities can be aligned with operational decision-making across transaction-intensive enterprises. The study adopts a structured narrative review of scholarly and technical literature on batch, micro-batch, stream-processing, Lambda, Kappa, event-driven, cloud-native, lakehouse and serverless architectures, with additional attention to governance, security, observability, resilience and emerging-economy implementation contexts. The findings indicate that no single architectural model is universally optimal. Batch processing remains valuable for reconciliation, regulatory reporting and historical analysis, whereas stream-oriented and event-driven designs are better suited to fraud detection, payment monitoring, inventory visibility and other latency-sensitive operations. Hybrid architectures offer the strongest balance between speed, correctness, recoverability and cost. The review further finds that scalability depends not only on distributed computing, but also on partitioning, state management, change data capture, automated testing, lineage, data contracts, quality controls and service-level objectives. Organisational alignment, cross-functional ownership and regulatory compliance are equally decisive in determining whether technical capability produces measurable business value. The study concludes that real-time analytical performance must be evaluated through both engineering and operational outcomes. It recommends use-case-driven architectural selection, resilient hybrid deployment, embedded security and governance, automated quality assurance, transparent AI-assisted pipeline management and stronger collaboration between technical and business teams. Future research should develop standardised benchmarks combining latency, accuracy, resilience, cost, sustainability and operational impact, while giving greater attention to infrastructure-constrained and emerging-economy environments. These priorities are essential for building data infrastructures capable of supporting responsive, evidence-based enterprise operations at scale.

View source

Similar papers

Open access Sep 2026

Serverless Data Engineering: Innovations in Python-Driven ETL Automation on AWS

Traditional cluster-based ETL architectures impose a structural tax on data engineering organisations: fixed compute resources provisioned for peak demand, scheduled batch cycles that introduce latency regardless of downstream urgency, and operational overhead that redirects engineering capacity from pipeline design to...

Rambabu Bolineni · 0 citations
Open access Aug 2026

Intelligent Business Integration With AI

An artificial intelligence-enhanced middleware pattern that augments existing integration stacks with telemetry, stream processing, and a lightweight learning loop to predict failures, automatically tune policies, and direct traffic in real time is presented.

Tejas Gajjar · 0 citations
Aug 2026

Reaching the Pinnacle of TPC-DS: Co-Design of Architecture, Executor, and Storage in TDSQL

Enterprise data explosion and the urgent industry demand for realtime complex multidimensional analytics require OLAP databases to be highly scalable, efficient, and cost-effective. Though diverse solutions (shared-nothing MPP databases, cloud-native decoupled systems, in-process analytical engines) exist with respecti...

Yi-Teng Chu, Jie Jiang, Yuxing Chen et al. · 0 citations
Open access Jul 2026

Self-Service Business Intelligence and Analytics in Digital Transformation

The digital-transformation era has changed how organizations generate, govern, and consume analytical insight. This paper analyzes the shift from centralized, IT-centric reporting toward self-service business intelligence (BI), in which business users author governed analyses without continuous engineering intervention...

Meena Jose Komban · 0 citations
Review Open access 2021

Real-Time Data Engineering for Financial Systems: Building Fault-Tolerant and High-Performance Pipelines

In today’s fast-paced financial industry, real-time data processing has become a crucial necessity for organizations to gain a competitive edge. Financial systems demand high-performance and fault-tolerant data pipelines capable of handling large-scale, high-velocity data streams while ensuring accuracy, security, and...

Shalini Gupta, D. Mishra · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.