Speaker Recognition Systems: A Review of Machine Learning and Deep Learning Approaches

Authors

  • Fadhliyatun Farah Mahadi Universiti Tun Hussein Onn Malaysia
  • Rafizah Mohd Hanifa Universiti Tun Hussein Onn Malaysia, Hab Pendidikan Tinggi Pagoh
  • Shamsul Mohamad Universiti Tun Hussein Onn Malaysia
  • Roszaini Haniffa Edinburgh Business School, Heriot-Watt University, SCOTLAND

Keywords:

Speaker recognition, machine learning, deep learning, MFCC, Wav2Vec, CNN, systematic review

Abstract

Speaker recognition has emerged as a cutting-edge technology with a diverse array of applications, ranging from biometric authentication to personalized user experiences. Its primary objective is to identify or authenticate individuals based on their unique vocal characteristics. This paper presents a comprehensive analysis of the evolution of speaker recognition systems, systematically comparing traditional statistical techniques with modern deep neural network–based approaches in terms of robustness, scalability, and computational efficiency. Historically, speaker recognition systems have relied on statistical models, including Gaussian Mixture Models (GMMs), Hidden Markov Models (HMMs), and Support Vector Machines (SVMs). While these methods provide a reliable performance under controlled conditions, they suffer notable degradation in the presence of speech variability, environmental noise, and large-scale datasets. In contrast, recent deep learning frameworks, including d-vector and x-vector architectures, have achieved substantial performance gains, with reported classification accuracies exceeding 99% in benchmark studies and relative improvements of 10–30% over traditional methods. Furthermore, speaker verification research increasingly emphasizes metrics such as Equal Error Rate (EER) and minimum Detection Cost Function (minDCF), with state-of-the-art systems reporting EER values as low as 3–4% and minDCF scores below 0.03 in challenging evaluation conditions. Key challenges, including spoofing attacks, dataset bias, multilingual variability, and privacy preservation, are critically examined. Based on the reviewed findings, this study concludes that deep learning–driven speaker recognition systems offer superior performance and adaptability for real-world deployment. Future research should focus on the ethical use of data, bias mitigation, and the development of lightweight yet secure models suitable for global and resource-constrained environments. 

Downloads

Download data is not yet available.

Downloads

Published

15-04-2026

Issue

Section

Special Issue 2026: ICon3E2025 (E)

How to Cite

Fadhliyatun Farah Mahadi, Rafizah Mohd Hanifa, Shamsul Mohamad, & Roszaini Haniffa. (2026). Speaker Recognition Systems: A Review of Machine Learning and Deep Learning Approaches. International Journal of Integrated Engineering, 18(1), 155-169. https://penerbit.uthm.edu.my/ojs/index.php/ijie/article/view/24379