Predicting CNN training latency on cloud GPUs without disclosing model architecture
This work presents PROFET, a system that predicts the training latency of arbitrary Convolutional Neural Network implementations across a wide range of GPU types and mini-batch sizes and introduces a novel operation-name clustering heuristic that effectively resolves naming inconsistencies in computational graphs where semantically similar operations are labeled differently.