Spatiotemporal vision transformers with Byzantine-robust federated prompt tuning for continuous urban perception
While spatial Vision Transformers (ViTs) achieve high precision in urban scene parsing, their frame-by-frame application in autonomous driving suffers from severe temporal flickering and prohibitive retraining costs across decentralized vehicle fleets. To overcome these dual bottlenecks, this paper introduce...