Raft-Consistent Metadata with GPU-Accelerated Encryption for Secure HDFS
Abstract
In today’s cloud computing environments, large-scale data infrastructures like Hadoop are facing severe challenges in providing efficient and secure static storage; these challenges stem from limitations in single algorithm encryption schemes as well as metadata management instability. This paper proposes a co-designed architecture demonstrating that metadata consistency and high-throughput encryption can be achieved simultaneously without performance trade-offs. Our approach integrates two components: (1) a Raft-consistent metadata service (RCMS) that replaces the single-point-of-failure NameNode with distributed consensus-based replication. Empirical evaluation shows that RCMS has a metadata fail-over latency of 1.35(+-0.42) seconds - an improvement of 78 times over ZooKeeper-based HA (105(+-28) seconds). Such fast recovery is essential for metadata-sensitive applications such as real-time analytics, streaming applications, and interactive SQL engines. Concomitantly, GHEM provides the encryption throughput of 935 MB/s, which is 51% uplift from CPU-based HDFS-TDE’s throughput of 620 MB/s. 8.1 times faster compared to 10 GB files. The key to achieving this architecture is that RCMS and GHEM are executed in non-overlapping critical paths and that metadata transactions and data encryption are performed in parallel with the network I/O, with a total overhead of less than 1%, given all together. This co-design confronts this entrenched notion that strong consistency forces performance compromises. Experimental validation for a four-node prototype cluster validates both the fast fail-over of metadata and the high-throughput encryption facilities. Finally, the proposed framework is easily extendable to various types of distributed storage systems, including Ceph, Cassandra, and cloud object storage, which makes this framework a solid foundation for building secure, consistent, and high-performance data storage solutions.