This work introduces a real-world multi-modal benchmark for adverse-weather autonomous driving, designed with three capability dimensions: observability awareness, spatial reliability, and risk-aware decision-making, enabling fine-grained diagnosis of model behavior under degraded observations.
Abstract
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains unclear how vision-language models behave under real-world adverse weather with multi-modal inputs. We argue that a key difficulty lies in degraded environmental observability: under fog, rain, snow, and low illumination, multi-modal observations become unreliable and cross-modally inconsistent, posing challenges to scene understanding, and subsequent decision-making. To study this, we introduce \textbf{ObsDriveBench}, a real-world multi-modal benchmark for adverse-weather autonomous driving. Our benchmark is designed with three capability dimensions: \textbf{observability awareness}, \textbf{spatial reliability}, and \textbf{risk-aware decision-making}, enabling fine-grained diagnosis of model behavior under degraded observations. We construct the benchmark through observability meta-annotation, scene description, and capability oriented multiple-choice tasks over synchronized camera, LiDAR, and radar inputs, forming a benchmark with over 14k training and 13k test questions. Experiments reveal consistent performance degradation of existing vision-language models. We further introduce \textbf{ObsDrive} model with normal-weather supervised fine-tuning and adverse-weather reinforcement learning, improving robustness across all three capabilities. The dataset and evaluation code will be released at \href{https://github.com/russellyq/ObsDriveBench}{\texttt{ObsDriveBench}}.
Standard deployment-ready object detectors for autonomous vehicles degrade in adverse weather and lighting conditions without being trained on extensive domain-specific data. While large-scale vision foundation models offer robust zero-shot generalization, their high computational cost makes them impractical for real-t...
Sepideh Gohari, Goodarz Mehr, A. Eskandarian· 0 citations
Reliable drone detection under real-world deployment conditions requires training data that spans the full operational design domain, including adverse weather and seasonal appearance variation. However, acquiring and annotating such data at scale remains highly resource-intensive, as adverse-weather conditions are inh...
T. Lenhard, Andreas Weinmann, Tobias Koch· 0 citations
With the rapid advancement of autonomous driving and Intelligent Transportation Systems (ITS), roadside perception—an essential component of vehicle-to-everything (V2X) communication—has become a critical foundation for large-scale traffic monitoring and data-driven safety decisions. However, under adverse environm...
Guo-Yu Zhang, Peng Hang, Xin Xia et al.· Communications in Transporta...· 0 citations
Autonomous vehicle perception systems rely heavily on robust semantic segmentation to interpret complex urban environments under dynamic driving conditions. While modern vision transformers (ViTs) demonstrate remarkable performance on pristine benchmarks, their accuracy degrades drastically in adverse weather condition...
Rashed Karim Bipul· International Journal of App...· 0 citations
For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high repeatability. Yet Multimodal Large Language Models (MLLMs) remain unreliable in continuous-valued measurement despite strong results on general multimodal benchmarks. Existing benchmarks expose this weakness but...
Chao-Wen Shen, Xin-Yuan Li, Yun Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.