Skip to content
Open access

Semantic Skyline: Multi-Embedding Skyline Query Processing for RAG Workloads in Vector Databases

2026 · IEEE Access · Vol 14, pp. 128780-128808 · 0 citations · 58 references

Abstract

Retrieval-Augmented Generation (RAG) systems and modern vector databases rely on embedding representations to retrieve semantically relevant entities. However, many real-world applications require multi-criteria query processing, where users express preferences across multiple semantic dimensions rather than a single notion of similarity. Existing approaches primarily rely on similarity-based nearest neighbor search or ranking-based aggregation, which either produce ambiguous results or incur significant preprocessing and maintenance overhead. Moreover, these approaches do not provide explicit support for multi-criteria dominance queries over embedding data. For example, in the real-world tourism service platform, it requires users to manually reason about trade-offs across embedding-based criteria (e.g., scenic beauty vs. urban convenience), thereby increasing user effort and platform load. In this work, we present a semantic skyline query processing framework for multi-embedding retrieval in RAG workloads. Our approach introduces a projection-based formulation of skyline dominance that captures directional semantic preferences through interpretable semantic axes. To support efficient execution, we design a hybrid semantic axis framework that combines reusable global structures with query-specific refinements, enabling effective candidate generation and pruning across diverse queries. In addition, we develop a workload-aware optimization strategy that determines when semantic structures are precomputed or computed on-the-fly, providing a practical trade-off between storage overhead and query execution cost. We evaluate the proposed framework on multi-embedding retrieval workloads across multiple datasets and embedding models. Experimental results show that the framework preserves directional semantic dominance while reducing query execution overhead through workload-aware semantic structure reuse in vector database-driven RAG systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.