Jul 2026· International Conference on Big Data Computing Service and Applications· pp. 125-129· 0 citations· 17 references
Abstract
Modern data platforms rely on pipeline-oriented architectures that are rigid, hard to adapt, and lack native auditability. We present Agentic Data Services, a control-planedriven architecture for Big Data as a Service (BDaaS) that models workflows as adaptive, policy-aware service entities rather than static directed acyclic graphs (DAGs). The architecture combines (i) a dual-record execution model adapted from pharmaceutical batch manufacturing-versioned Master Batch Records (MBRs) for workflow definition and immutable Electronic Batch Records (EBRs) for execution traces-and (ii) workflow-level semantic caching that reuses results across semantically similar requests. We implement the system as Agentic DataHub, a set of Rustbased microservices deployed on Kubernetes, and evaluate the semantic-caching component on a reproducible benchmark using sentence-transformer embeddings (all-MiniLM-L6-v2) and a FAISS flat inner-product index. For clustered workloads-semantically related requests grouped into 10 clusters with 70% intra-cluster similarity-the cache reduces backend requests by 81% and median latency by 93%, with 40% P95 latency reduction. We discuss generalization across domains and the architectural constraints that bound these results.
This work presents DeepEye, a workflow-centric agentic data system that turns user intents into transparent and steerable analytical workflows and develops DataMagic as the system’s Video Generator, a declarative multi-agent method that improves data-video quality.
This work presents DeepEye, a workflow-centric agentic data system that turns user intents into transparent and steerable analytical workflows and develops DataMagic as the system’s Video Generator, a declarative multi-agent method that improves data-video quality.
VexDB is presented, an agent-ready, cross-platform, multi-modal database system designed to unify data management across on-device, on-premises, and cloud-native deployments, and decouples the data architecture from the infrastructure architecture, allowing the same system abstractions to operate over diverse compute s...
Traditional cluster-based ETL architectures impose a structural tax on data engineering organisations: fixed compute resources provisioned for peak demand, scheduled batch cycles that introduce latency regardless of downstream urgency, and operational overhead that redirects engineering capacity from pipeline design to...
Rambabu Bolineni· East African Journal of Info...· 0 citations
DataClawEval is introduced, the first comprehensive benchmark designed specifically to evaluate the end-to-end task completion capabilities of autonomous agents in real-world data engineering scenarios, and it comprises 100 rigorous, end-to-end tasks spanning five execution engines.
Cloud native data pipelines have become an enabling ingredient of the modern enterprise analytics to fulfill the ever-increasing demand of a scale-loving, resilient, and real-time processing of a wide range of data sources. Organizations currently produce large amounts of structured, semi-structured, and unstructured d...
Ethan Williams· International Journal of App...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.