PruneInfer: Exact Full-Neighborhood GNN Inference of Large Datasets on a Single GPU via Full Topology Pruning
Graph neural network (GNN) inference in deployment often requires deterministic and exact predictions, which in turn require each inference run to aggregate complete dependency information from the full graph topology. However, over large graphs, full-graph forward propagation is usually infeasible on GPU due to limite...