Spatio-Temporal Masked Autoencoders with Contrastive Cross-Attention for Robust Video Semantic Segmentation in Adverse Autonomous Driving Conditions
Autonomous vehicle perception systems rely heavily on robust semantic segmentation to interpret complex urban environments under dynamic driving conditions. While modern vision transformers (ViTs) demonstrate remarkable performance on pristine benchmarks, their accuracy degrades drastically in adverse weather condition...