Development of a Deep Ensemble Framework for Speech based Authentication System

Arpita Choudhury, Pinki Roy, Sivaji Bandyopadhyay

Abstract


Speech recognition has become ubiquitous on numerous platforms, such as AI-driven applications. Despite its extensive use, speech-based authentication remains less explored compared to traditional text-based methods. Every individual’s voice is unique due to inherent anatomical differences. This uniqueness can be effectively employed in authentication applications by extracting and modeling these discriminative vocal features. Moreover, speech-based authentication is advantageous for hands-free operations and accessibility for individuals with disabilities. Integrating speech into a multi-factor authentication framework can significantly enhance security, making it more resilient against real-time fraudulent attempts. Nevertheless, achieving high speech recognition accuracy in the presence of intense background noise remains a major challenge. This work proposes a deep ensemble framework for speech recognition. The vVISWa and AudioMNIST datasets are used for experimentation. Both datasets are pre-processed to obtain audio samples of single-digit and digit sequence (6-digit) utterances. The audio samples are further augmented to mimic real-world conditions, such as extreme background noise and variation in speed and volume across speakers. The ensemble comprises three lightweight convolutional neural network (CNN) architectures: MobileNet, EfficientNetB0, and DenseNet121. The features extracted by these networks from mel-spectrogram representations of audio signals are integrated and fed into a squeeze-and-excitation block for recalibration. A fully connected softmax layer uses the recalibrated feature to perform classification. Under three-fold cross-validation, each model demonstrated significant performance gain. The average accuracy achieved by the ensemble model exceeds 95\% for both the original(i.e. clean) and the augmented datasets. This makes the proposed system well-suited for deployment in secure, speech-driven authentication applications.

Keywords


Speech recognition, ensemble, deep learning, authentication.

Full Text: PDF