Skip to content

VoiceNet: A Responsible Artificial Intelligence Multilingual Speech Framework for Gender and Age Classification.

Sep 2026 · Journal of Visualized Experiments · Vol 235 · 0 citations
Medicine

Abstract

Accurate gender and age classification from voice data is important for personalized, secure, and ethical human-computer interaction (HCI). With the growing use of large pretrained speech models such as wav2vec 2.0, there is a need to leverage these representations responsibly for multilingual and privacy-aware demographic inference. This study proposes VoiceNet-RAI, a deep learning (DL) framework that integrates Mel-Frequency Cepstral Coefficients (MFCCs) with transformer-based multilingual wav2vec 2.0 embeddings within an EfficientNet-Lite architecture to support accurate, fair, and transparent classification across diverse linguistic contexts. The framework combines contextual and spectral information to improve robustness under variable recording conditions. Experiments on multilingual speech data covering 28 languages demonstrate strong classification performance, with the English subset achieving 97.87% accuracy, 98.61% precision, 98.70% recall, and a 98.65% F1-score. External validation using the Mozilla Common Voice dataset yielded 98.27% accuracy for gender classification, demonstrating generalizability to an independent dataset. Ablation analysis further showed that combining MFCC and wav2vec 2.0 features improved performance compared with raw and MFCC-only representations. These findings demonstrate the potential of hybrid spectral-contextual speech representations for robust and responsible gender and age classification.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.