Open-Vocabulary Forgetting in the Downstream Adaptation of Remote Sensing Vision–Language Detectors
Abstract
Open-vocabulary detectors pretrained on remote sensing imagery can be queried with free-form category names instead of a fixed taxonomy. In practice, they reach applications through fine-tuning on small, specialized target datasets, and standard evaluation reports only target-dataset accuracy, so any loss of the pretrained open-vocabulary capability goes unmeasured. In this article, we present a retention-aware adaptation benchmark for remote sensing vision–language detectors: a fixed dual-axis protocol that evaluates every adapted model on target accuracy and on open-vocabulary retention, measured as mean average precision (mAP) on the 80-class LAE-80 C benchmark, with retention never used for model selection. Applying the protocol to LAE-DINO across one in-domain and three held-out targets shows that in-domain fine-tuning reduces retention from the pretrained 24.1 mAP by 0.8 mAP (3%), whereas held-out fine-tuning reduces it by 5.6 to 8.6 mAP (23% to 36%); under the fixed optimization budget of our experiments, the tested few-shot configurations lost more retention than full-data adaptation. Classwise and component-level diagnostics locate the loss on classes semantically close to the target vocabulary and show that it is not recoverable by restoring any single tested component, with prompt masking attributing about 12% of the distant-class loss to score competition. Across seven mitigation strategies, none produces a practically meaningful retention recovery at matched target accuracy or meets the predefined success criterion; weight interpolation of the final model dominates the early checkpoints, matches the later ones within the test-set uncertainty, and defines the best observed accuracy–retention frontier among the tested configurations. A second detector configuration, evaluated on its own natural-image vocabulary, reproduces the direction of these regularities. We release the protocol, splits, prediction files, and checkpoints as a reproducible benchmark.