Skip to content
Open access

AI4TEN: Fine-Tuning Pre-Trained Audio Transformers (BEATs, AST) for Cross-Dataset Acoustic Vehicle Classification with Domain Adaptation

Sep 2026 · Acoustics · Vol 8, pp. 63 · 0 citations · 12 references

TL;DR

Both pre-trained audio transformers substantially outperform the baseline for cross-dataset vehicle classification into five categories, with self-supervised pre-training producing slightly more transferable representations than supervised pre-training.

Abstract

Acoustic vehicle classification from roadside microphones supports source-specific traffic noise monitoring, but classifiers trained on one dataset generalize poorly to recordings from different environments. This study compares pre-trained audio transformers (BEATs, AST) against a CNN baseline for cross-dataset vehicle classification into five categories (car, truck, motorcycle, bus, background), training on the IDMT-Traffic dataset (Germany) and testing on the MELAUDIS dataset (Australia). Two domain adaptation methods (DANN, ArcFace) are applied to the best-performing transformer. BEATs achieves a balanced F1 of 0.46, a 77% improvement over the CNN baseline (0.26). AST achieves F1 = 0.43, a 65% improvement. Both pre-trained transformers substantially outperform the baseline, with self-supervised pre-training producing slightly more transferable representations than supervised pre-training. ArcFace metric learning achieves the highest cross-dataset F1 (0.50), modestly outperforming standard fine-tuning. DANN degrades cross-dataset performance but achieves the highest microphone robustness score (F1 = 0.80). Classification of the underrepresented bus class (53 training samples) improves from F1 = 0.38 to 0.54 with ArcFace. Microphone robustness is strong across all models (BEATs F1 = 0.75), confirming that recording environment dominates the domain shift over hardware variation. The choice of domain adaptation strategy should be guided by the expected type of domain shift.

Read PDF

Similar papers

Sep 2026

Bridging the Synthetic-to-Real Gap in Impulsive Sound Detection Using Audio Transfer Learning

A CNN achieving 98.4% F1 on synthetic benchmark spectrogram data collapses to 20.0% F1 on real-world audio, revealing a severe synthetic-to-real domain gap in impulsive sound detection. This paper provides one of the first quantitative studies of this phenomenon and demonstrates that representation choice dominates cla...

Charlie Holden, Mathew Ridgely, Jayanth Bhansali et al. · 0 citations
Open access Oct 2026

Fusion-Inception-Aux: An Auxiliary-Supervised Multi-Scale Fusion Network for Distributed Acoustic Sensing Event Recognition

Results indicate that multi-scale Inception encoding and auxiliary supervision are the most important components, while the PatchX interaction module provides no consistent additional gain under the current setting.

Tian-Chang Xie, Hai-Ling Wang, Wei-Guang Wang et al. · 0 citations
Preprint Aug 2026

Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

PRISM (Prototype-Rectified Iterative Self-supervised Manifold Denoising), a training-free, source-free TTA framework grounded in the Affine Noise Hypothesis, is addressed, making it substantially faster than gradient-based TTA while requiring no additional training.

A. Shukla, R. Thakur, Aryan Das et al. · 0 citations
Open access Sep 2026

Domain-specific unsupervised pre-training for robust respiratory sound classification

This framework provides a potential scalable pathway for domains with critically limited annotated data using massive synthetic data without synthetic labels to train a lightweight classifier on limited clinical data.

Takehiro Hirasawa, Yasumasa Tamura, K. Shimizu et al. · 0 citations
Preprint Aug 2026

SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification

SPECTRA is a framework built on a frozen encoder which adds three components, a lightweight trainable adapter that calibrates the generic embeddings to the task, an exemplar-free anti-forgetting scheme, and a transductive optimal-transport refinement of prototypes at test time.

Giries Abu Ayoub, L. Mualem, Simon Korman · 0 citations
Preprint Aug 2026

U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement

U-PAST is a hybrid transformer-U-Net architecture that addresses self-attention dependency-modeling in the complex spectrogram domain through self-attention dependency-modeling in the complex spectrogram domain, offering an attractive performance-to-cost trade-off at a small parameter footprint.

Cao Duong Ly, J. Anemüller · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.