Machine Learning-Based Data Sampling Techniques for Big Data Analytics
Abstract
The rapid expansion of IoT, cloud computing, social media, and distributed systems has generated massive amounts of data, making big data analytics increasingly important. Traditional sampling methods are often insufficient for handling the complexity and scalability challenges of modern big data environments. To overcome these limitations, machine learning-based sampling techniques have been developed to intelligently select representative data while preserving analytical accuracy. This paper surveys supervised, unsupervised, reinforcement, active, and deep learning-based sampling approaches and their integration with platforms like Hadoop and Apache Spark. Experimental findings show that intelligent sampling improves scalability, accuracy, efficiency, and resource utilization. The study concludes that machine learning-based sampling is a key technology for future big data analytics, with promising directions in hybrid, federated, explainable, and real-time adaptive sampling models.