Skip to content
Review Open access

Data-Centric Security for Generative AI Systems: Protecting Sensitive Data against Leakage, Inference Attacks, and Unauthorized Model Access

2026 · International Journal of Data Engineering and Intelligent Computing · Vol 9, pp. 1-32 · 0 citations

TL;DR

A data-centric security framework is proposed that links data sensitivity, access policies, model controls, and governance requirements throughout the AI lifecycle and emphasizes that security controls should remain tied to sensitive information regardless of where the data is stored, processed, retrieved, or generated.

Abstract

The growing use of generative artificial intelligence in enterprise and public-sector environments has introduced significant concerns regarding the protection of sensitive data. Generative AI systems process large volumes of information across training, fine-tuning, retrieval, inference, and output stages, creating multiple opportunities for unauthorized disclosure and misuse. This study examines data-centric security as an approach for protecting sensitive information against data leakage, membership inference, model inversion, prompt-based attacks, retrieval-related exposure, model extraction, and unauthorized access to AI services. It reviews major attack surfaces across the generative AI lifecycle and evaluates security controls including data classification, minimization, encryption, privacy-preserving learning, identity and access management, secure retrieval-augmented generation, output filtering, data loss prevention, and continuous monitoring. Based on these findings, the study proposes a data-centric security framework that links data sensitivity, access policies, model controls, and governance requirements throughout the AI lifecycle. The framework emphasizes that security controls should remain tied to sensitive information regardless of where the data is stored, processed, retrieved, or generated. The study provides practical guidance for organizations seeking to reduce privacy and confidentiality risks while maintaining controlled and accountable use of generative AI systems.

Read PDF

Similar papers

Review Open access Sep 2026

Bridging ML and Threat Landscapes: Inference Attacks and Cryptographic Defenses for Training-Phase Data Protection

Machine learning systems use sensitive personal data to train and infer models, but developers are largely unaware of privacy attack vectors that evade conventional data security controls. This paper surveys the privacy threat landscape inherent to ML workflows, including membership inference attacks, model inversion a...

Pramod Prakash · 0 citations
Review Sep 2026

Privacy-preserving methodologies against privacy attacks on deep learning: a survey

The analysis finds that noise-based methods such as DP remain the practical baseline but offer only partial protection; cryptographic approaches provide stronger theoretical guarantees at substantially higher cost; and LLM/multimodal leakage remains an urgent, under-benchmarked gap.

Subhasish Ghosh, A. K. Mandal · 0 citations
Review Open access Aug 2025

Data security in large language models: risks, defense, and directions

Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine translation, and conversational systems. Despite their transformative potential, these models inherently rely on massive amounts of training data, often collected from diverse and...

Kang Chen, Xiu-Ze Zhou, Yuan-Hui Yu et al. · 2 citations
Conference Aug 2026

Protecting Intelligent Information Systems: Security Threats, Privacy Risks, and Defense Mechanisms

Intelligent Information Systems (IIS) have taken center stage in the modern digital infrastructures and have allowed sophisticated decision-making in healthcare, finance, transportation, smart cities, and in cyber-physical systems. Their increasing reliance on big data, machine learning models, and distributed systems,...

Shyni Shajahan, S. Swetha, J. Elavarasi et al. · 0 citations
#federated learning Open access Sep 2026

Research on Copyright Governance and Security Challenges of Training Data for Generative Artificial Intelligence

Generative AI depends on training with very large volumes of data, and the acquisition and use of that data has become a focal point for copyright infringement and for privacy and security risks. This paper examines how generative AI training data is governed and what security problems it raises. It compares the reason...

Hui-Yao Jian · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.