PipeTree: Decision Tree Partitioning for Efficient Inference on Programmable Switches
Abstract
In recent years, there has been a growing trend in deploying machine learning models directly in the data plane, taking advantage of the high throughput and low latency offered by modern programmable switches to conduct various tasks such as line-rate traffic classification and anomaly detection. Among these models, decision tree has gained popularity due to its simple structure and top-down matching pattern, which aligns well with the typical pipeline architecture of programmable switches. However, deploying decision tree on pipeline-based programmable switches, such as Tofino, is challenging due to uneven resource utilization, together with the depth and width limitations imposed by the pipeline resources. In this paper, we propose a novel tree partitioning approach called PipeTree that enables parallel execution of decision tree through making full use of the pipeline resources. Specifically, PipeTree cuts the deep decision tree into multiple short subtrees which could be executed in parallel and uses a composition table to aggregate the outputs of all subtrees. The experimental results on programmable switches show that, compared with state-of-the-art data plane decision tree approaches, PipeTree enhances the scalability of decision tree and effectively reduces the inference time.