Benchmarking Test-Time Adaptation for Multi-Label Chest X-ray Classification under Distribution Shift
Abstract
Although AI models achieve impressive performance on chest X-ray benchmarks, their deployment in real-world clinical settings remains challenging due to performance degradation under unpredictable conditions. To understand how these AI models can adapt in practice, we systematically evaluate several Test-Time Adaptation (TTA) techniques for multi-label classification under two key challenges: the natural domain shifts when moving to a new hospital, and non-stationary environments involving image corruptions that evolve either gradually over time or as abrupt shifts. Our experiments on the CheXpert and NIH-14 datasets reveal distinct strengths across different scenarios: RoTTA, with its long-term memory, excels in gradually changing environments, achieving a 5.37% improvement in mean AUC compared to its zero-shot baseline. In contrast, CoTTA, with its stochastic restoration mechanism, emerges as the best choice for handling abrupt shifts in data by achieving a stable 0.62% improvement. These results highlight the fundamental trade-off in TTA between maintaining long-term consistency and enabling immediate responsiveness. However, a significant open challenge remains in optimizing memory refreshment logic to resolve the immortal sample bottleneck, ensuring sustained resilience in highly volatile environments.