Blind Device-Response Interpolation for Intermediate-Domain Unsupervised Adaptation in Cross-Device Acoustic Scene Classification
Abstract
Recording-device mismatch is one of the dominant failure modes of acoustic scene classification: a model trained on one microphone degrades sharply on an unseen one. Aligning source and target in a single adversarial step is unstable when the device gap is large, and recent intermediate-domain formulations mitigate this by inserting bridge domains, but they build those bridges from hand-specified parametric device operators that are unavailable at deployment time. We propose BRIDA, an intermediate-domain unsupervised adaptation framework whose bridges are derived from the device response itself. A blind estimator recovers the relative log-mel magnitude response between the labelled source devices and the unlabelled target device from long-term spectral statistics alone, with an optional class-balanced refinement for the case where the two unlabelled pools differ in scene composition. Because a magnitude filter is additive in the log-mel domain, fractional powers of the estimate yield a continuum of label-preserving bridge domains at essentially zero cost; we traverse that continuum with a curriculum and regularise it with an ordinal path-position adversary, a residual maximum-mean-discrepancy anchor and confidence-gated target prototypes. Experiments use 1648 real recordings and 66 measured microphone impulse responses under a leave-one-device-out protocol with disjoint source, target and evaluation content. Averaged over 3 unseen devices and 3 seeds, BRIDA attains 53.00 ± 4.29% target accuracy and 52.99 ± 4.58% macro-F1, improving on source-only training by 7.83 points and on a discrete two-bridge intermediate-domain baseline by 0.61 points, with no inference-time overhead.