Online Reinforcement Learning From Existing Heuristics for TCP Congestion Control
Abstract
Avoiding network bottlenecks while maximizing network utilization is paramount to supporting modern networked applications. Traditional congestion control protocols like Cubic, BBR, and TCP Vegas have demonstrated efficacy within specific network conditions. However, years of research on transport protocols have shown that no protocol works for all scenarios. To tame congestion, new machine learning (ML) models were proposed with promising results, though relying on extensive training and specific network configurations. To mitigate some of these adaptability challenges, we propose Mutant, an online reinforcement learning algorithm for congestion control that adapts to the behavior of the best-performing schemes, outperforming them in most network conditions. The key intuition behind our protocol design is that we should not merely learn from past values of the congestion window; instead, we should learn from the behavior of existing protocols since some of them behave very well in given scenarios but poorly in others. Design challenges included determining the best protocols to learn from in a given network scenario and developing a system that can adapt to future protocols with minimal changes. Our evaluation in real and emulated scenarios shows that Mutant achieves lower delays and higher throughput than previous learning-based systems while remaining fair by exhibiting negligible harm to competing flows, making it robust under diverse and dynamic network conditions.