Skip to content
Open access

Selecting the Right Data Pipeline Architecture for Reliability and Scale

2025 · International Journal of Data Engineering and Intelligent Computing · Vol 8, pp. 01-16 · 0 citations

TL;DR

A formal model that assists practitioners and designers to construct resilient, adaptable, and budget friendly data pipelines, which are in line with the current data needs is delivered.

Abstract

Today's data-driven systems lean quite a bit on the smooth operation of data pipeline architectures to manage the intake, processing, and redistribution of data at large volumes, thus emphasizing reliability and scalability as key aspects of design. As companies rely more on immediate information and extensive data analytics, the choice of pipeline architecture-going for batch, streaming, or hybrid-is a significant decision. Batch processing delivers straightforwardness and is a low-cost solution to handle periodic workloads while streaming set-ups provide data in real time with minimal delay. Hybrid methods try to bridge the gap by allowing both real time and historical cases. Nevertheless, identifying the right architecture means going through some major problems such as latency limits, fault tolerance, increased complexity in operations, and cost reduction. This paper lays out a clear decision making process for matching your data pipeline architecture to your workload patterns, system specs, and scale factors. Besides including the basics of good design, the use of technologies and performance figures, the one-stop guide is also there to help you make wise decisions. A case study from the field is also examined to shed light on tool use, decisions, and side-effects through real work situations. The results reveal that one architecture is not capable of covering all use cases; instead, working based on the conditions is the winning strategy. Besides that, this paper delivers a formal model that assists practitioners and designers to construct resilient, adaptable, and budget friendly data pipelines, which are in line with the current data needs.

Read PDF

Similar papers

Preprint Sep 2026

NDS: Programmer-Free Offload of High-Performance Near Data Strands

Near Data Processing (NDP) has the potential to significantly improve system performance and energy by alleviating data movement bottlenecks. However, most NDP proposals pose heavy requirements for the software stack, data layout, and/or the underlying hardware. To broaden NDP adoption, this work focuses on a modular h...

Shreyas Singh, Pratyush Nandi, Lin Jia et al. · 0 citations
Open access Aug 2026

HBA: A Heuristic Batch Advisor for Self-Tuning Bulk Data Ingestion in the OutSystems Low-Code Platform

Low code development platforms like OutSystems allow teams to build data-centric applications using visual models that the platform takes care of the database plumbing [1] [2]. We have demonstrated in an earlier study that the write path is row-wise and fails at scale; and in a previous study we introduced EAVS [25], a...

Anamika Garg, Sachin H. Patel · 0 citations

T 𝒆𝒙𝑩𝒆𝒏𝒄𝒉 : A Unified Benchmarking Suite for Shifting Workloads

A unified key-value benchmarking suite that enables benchmarking key-value stores against dynamically shifting and production-like workloads and comparing their performance side by side and allows users to benchmark multiple databases and perform an apples-to-apples comparison under the same workload readily within a s...

Abhishek Chanda, Shubham Kaushik, A. Lavrov et al. · 0 citations
Preprint Aug 2026

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.

Hanzhao Wang, Jingxuan Wu, Yumeng Li et al. · 0 citations
Preprint Aug 2026

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.

Xingqi Cui, Chieh-Jan Mike Liang, Ziang T. Tang et al. · 0 citations
Review 2025

Scalable ETL Pipeline Architectures for Real-Time Transaction Analytics: Bridging Data Engineering and Business Operations

The study concludes that real-time analytical performance must be evaluated through both engineering and operational outcomes, and recommends use-case-driven architectural selection, resilient hybrid deployment, embedded security and governance, automated quality assurance, transparent AI-assisted pipeline management a...

Ogochukwu T. Izuchukwu, Dominic Feboh, Ayokunle Olamide Ijagbemi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.