Integrating 2PC with Consensus for Fast Replication
Abstract
Fast consensus protocols, such as CURP-Q and NOPaxos, have coupled objectives of reducing latency and tolerating faults, but doing so incurs considerable processing overhead and/or necessitates complex changes to existing architectures. Our insight is that when no adverse execution conditions (including operation conflicts, node failures, and network failures) occur, replication can be done safely without consensus at all. Rather than pursuing one single consensus protocol for fast and fault-tolerant replication, we propose to decouple low latency from fault tolerance, and design a hybrid solution that seamlessly integrates client-coordinated two-phase commit (2PC) with consensus, respectively to reduce latency in normal situations and tolerate faults in faulty situations. In the absence of faults, the system enters the Fast mode, where the client directly broadcasts its requests to all replicas with 2PC completing all updates in the first phase, i.e., in one RTT. Otherwise, the system enters the Consensus mode, which resorts to a consensus protocol to tolerate faults. We have applied the hybrid solution to the widely-used Raft consensus protocol to design xRaft, which can adaptively switch between the Fast and Consensus modes, ensuring correctness with minimal switching overhead. Evaluation shows that xRaft significantly outperforms state-of-the-art consensus protocols.