Both pre-trained audio transformers substantially outperform the baseline for cross-dataset vehicle classification into five categories, with self-supervised pre-training producing slightly more transferable representations than supervised pre-training.
Abstract
Acoustic vehicle classification from roadside microphones supports source-specific traffic noise monitoring, but classifiers trained on one dataset generalize poorly to recordings from different environments. This study compares pre-trained audio transformers (BEATs, AST) against a CNN baseline for cross-dataset vehicle classification into five categories (car, truck, motorcycle, bus, background), training on the IDMT-Traffic dataset (Germany) and testing on the MELAUDIS dataset (Australia). Two domain adaptation methods (DANN, ArcFace) are applied to the best-performing transformer. BEATs achieves a balanced F1 of 0.46, a 77% improvement over the CNN baseline (0.26). AST achieves F1 = 0.43, a 65% improvement. Both pre-trained transformers substantially outperform the baseline, with self-supervised pre-training producing slightly more transferable representations than supervised pre-training. ArcFace metric learning achieves the highest cross-dataset F1 (0.50), modestly outperforming standard fine-tuning. DANN degrades cross-dataset performance but achieves the highest microphone robustness score (F1 = 0.80). Classification of the underrepresented bus class (53 training samples) improves from F1 = 0.38 to 0.54 with ArcFace. Microphone robustness is strong across all models (BEATs F1 = 0.75), confirming that recording environment dominates the domain shift over hardware variation. The choice of domain adaptation strategy should be guided by the expected type of domain shift.
A CNN achieving 98.4% F1 on synthetic benchmark spectrogram data collapses to 20.0% F1 on real-world audio, revealing a severe synthetic-to-real domain gap in impulsive sound detection. This paper provides one of the first quantitative studies of this phenomenon and demonstrates that representation choice dominates cla...
Charlie Holden, Mathew Ridgely, Jayanth Bhansali et al.· International Symposium on N...· 0 citations
Results indicate that multi-scale Inception encoding and auxiliary supervision are the most important components, while the PatchX interaction module provides no consistent additional gain under the current setting.
Tian-Chang Xie, Hai-Ling Wang, Wei-Guang Wang et al.· IEEE Photonics Journal· 0 citations
PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.
A. Shukla, R. Thakur, Aryan Das et al.· 0 citations
This framework provides a potential scalable pathway for domains with critically limited annotated data using massive synthetic data without synthetic labels to train a lightweight classifier on limited clinical data.
Takehiro Hirasawa, Yasumasa Tamura, K. Shimizu et al.· Artificial Life and Robotics· 0 citations
SPECTRA is a framework built on a frozen encoder which adds three components, a lightweight trainable adapter that calibrates the generic embeddings to the task, an exemplar-free anti-forgetting scheme, and a transductive optimal-transport refinement of prototypes at test time.
Giries Abu Ayoub, L. Mualem, Simon Korman· 0 citations
U-PAST is a hybrid transformer-U-Net architecture that addresses self-attention dependency-modeling in the complex spectrogram domain through self-attention dependency-modeling in the complex spectrogram domain, offering an attractive performance-to-cost trade-off at a small parameter footprint.
Cao Duong Ly, J. Anemüller· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.