Enhancing speaker verification performance in noisy environments via parallel model combination

Authors

  • Duraid Y. Mohammed School of education for women, Al-Iraqia University, Baghdad, Iraq Author
  • Ahmed Aljuboori College of Education for Pure Science Ibn-Al-Haitham, University of Baghdad, Baghdad, Iraq Author

Keywords:

speaker verification, HMM, robustness, PMC, noise

Abstract

In real-world speaker applications, background and additive and convoluted noise and reverberation cause a mismatch between the training and recognition environments, resulting in poor performance. Parallel Model Combination (PMC) has proven to be an effective method for mitigating the effects of additive noise. Hidden Markov Model (HMM)-based speaker recognizers have been successfully improved using Parallel Model Combination (PMC). The results of using PMC to compensate for additive noise in HMM-based text-independent speaker verification are presented in this paper. Speech data were obtained from the SALU-AC databases. The findings of using PMC to correct for additive noise in HMM-based text-independent speaker verification are presented in this article. Equal Error Rates (EER) for noise-corrupted speech are given for various signal-to-noise ratios (SNRs) and noise sources. For example, at 10dB SNR, the average EER for speech in an operations room noise decreased from 6.7% uncompensated to less than 4.6% using PMC. Finally, it is shown that the precise value of the parameter that controls the relative amplitudes of the PMC model's speech and noise components does not affect speaker verification accuracy. Speaker verification experiments in artificial conditions show the efficiency of using Parallel Model Combination (PMC) in reducing equivalent EER error rate and DET trade-off error detection.

Downloads

Download data is not yet available.

References

[1] K. A. Al-Karawi, A. H. Al-Noori, F. F. Li, and T. Ritchings, "Automatic Speaker Recognition System in Adverse Conditions--Implication of Noise and Reverberation on System Performance," International Journal of Information and Electronics Engineering, vol. 5, pp. 423-427, 2015.

[2] M. M. Abdelwahab, K. A. Al-Karawi, E. Hasanin, and H. Semary, "Autism spectrum disorder prediction in children using machine learning," Journal of Disability Research, vol. 3, pp. 1–9, 2024.

[3] J. P. Campbell, "Speaker recognition: a tutorial," Proceedings of the IEEE, vol. 85, pp. 1437-1462, 1997.

[4] K. A. Al-Karawi and D. Y. Mohammed, "Improving short utterance speaker verification by combining MFCC and Entrocy in Noisy conditions," Multimedia Tools and Applications, vol. 80, pp. 22231-22249, 2021.

[5] K. A. Al-Karawi and D. Y. Mohammed, "Using combined features to improve speaker verification in the face of limited reverberant data," International Journal of Speech Technology, vol. 26, pp. 789-799, 2023.

[6] K. A. Al-Karawi and F. Li, "Robust speaker verification in reverberant conditions using estimated acoustic parameters—A maximum likelihood estimation and training on the fly approach," in 2017 Seventh International Conference on Innovative Computing Technology (INTECH), 2017, pp. 52-57.

[7] K. A. Al-Karawi and S. T. Ahmed, "Model selection toward robustness speaker verification in reverberant conditions," Multimedia Tools and Applications, vol. 80, pp. 36549–36566, 2021.

[8] K. A. Al-Karawi, "Face mask effects on speaker verification performance in the presence of noise," Multimedia Tools and Applications, vol. 83, pp. 4811-4824, 2024.

[9] N. Alotaibi, I. Al-Dayel, I. Elbatal, M. R. Ezzeldin, M. Elgarhy, and K. A. Al-karawi, "STATISTICAL ANALYSIS OF COVID-19 PANDEMIC IN SAUDI ARABIA," Advances and Applications in Statistics, vol. 74, pp. 107-118, 2022.

[10] K. A. Yousif, "Characters and Digits Recognition Using Neural Network Learned by Particle Swarm Optimization," Diyala Journal For Pure Science, vol. 10 pp. 73-91, 2014.

[11] K. A. Al-Karawi, "Mitigate the reverberation effect on the speaker verification performance using different methods," International Journal of Speech Technology, vol. 24, pp. 143-153, 2021/03/01 2021.

[12] A. Aljuboori, L. A. Tawfeeq, and K. A. Al-Karawi, "Pushing towards ehealth for iraqi hypertensives: an integrated class association rules into SECI model," Indonesian Journal of Electrical Engineering and Computer Science, vol. 22, pp. 522-533, 2021.

[13] M. M. Abdelwahab, K. A. Al-Karawi, and H. E. Semary, "Deep learning-based prediction of Alzheimer’s disease using microarray gene expression data," Biomedicines, vol. 11, p. 3304, 2023.

[14] D. A. Reynolds, "Speaker identification and verification using Gaussian mixture speaker models," Speech communication, vol. 17, pp. 91-108, 1995.

[15] D. A. Reynolds, T. F. Quatieri, and R. B. Dunn, "Speaker verification using adapted Gaussian mixture models," Digital signal processing, vol. 10, pp. 19-41, 2000.

[16] A. S. Alenizi and K. A. Al-karawi, "Cloud Computing Adoption-Based Digital Open Government Services: Challenges and Barriers," in Proceedings of Sixth International Congress on Information and Communication Technology, 2022, pp. 149-160.

[17] K. A. Al-Karawi and D. Y. Mohammed, "Early reflection detection using autocorrelation to improve robustness of speaker verification in reverberant conditions," International Journal of Speech Technology, vol. 22, pp. 1077–1084, 2019.

[18] K. A. Al-Karawi and B. Al-Bayati, "The effects of distance and reverberation time on speaker recognition performance," International Journal of Information Technology, vol. 16, pp. 3065-3071, 2024.

[19] J. Ortega-García and J. González-Rodríguez, "Overview of speech enhancement techniques for automatic speaker recognition," in Spoken Language, 1996. ICSLP 96. Proceedings., Fourth International Conference on, 1996, pp. 929-932.

[20] A. S. Alenizi and K. A. Al-karawi, "Machine learning approach for diabetes prediction," in International Congress on Information and Communication Technology, 2023, pp. 745-756.

[21] A. P. Varga and R. K. Moore, "Hidden Markov model decomposition of speech and noise," in International Conference on Acoustics, Speech, and Signal Processing, 1990, pp. 845-848.

[22] J. Ming, T. J. Hazen, J. R. Glass, and D. A. Reynolds, "Robust speaker recognition in noisy conditions," Audio, Speech, and Language Processing, IEEE Transactions on, vol. 15, pp. 1711-1723, 2007.

[23] M. J. Gales and S. J. Young, "Robust continuous speech recognition using parallel model combination," IEEE Transactions on Speech and Audio Processing, vol. 4, pp. 352-359, 1996.

[24] A. H. Al-Noori, K. A. Al-Karawi, and F. F. Li, "Improving Robustness of Speaker Recognition in Noisy and Reverberant Conditions via Training," in Intelligence and Security Informatics Conference (EISIC), 2015 European, 2015, pp. 180-180.

[25] M. M. Abdelwahab, K. A. Al-Karawi, and H. Semary, "Integrating gene selection and deep learning for enhanced Autisms' disease prediction: a comparative study using microarray data," AIMS Mathematics, vol. 9, pp. 17827–17846, 2024.

[26] K. A. Al-karawi, "Internet of Things (IoT) about Disabilities: Disabilities in relation to the Internet of Things (IoT)," ScienceOpen Preprints, 2023.

[27] A. S. Alenizi and K. A. Al-Karawi, "Internet of things (IoT) adoption: challenges and barriers," in Proceedings of Seventh International Congress on Information and Communication Technology: ICICT 2022, London, Volume 3, 2023, pp. 217-229.

[28] d. Mohammed, K. A. Al-Karawi, P. Duncan, and F. F. Li, "Overlapped Music segmentation using a new Effective Feature and Random Forests," International Journal Of artificial intelligence (IN-IA), vol. 8, pp. 181-189, June 2019.

[29] A. S. Alenizi and K. A. Al-Karawi, "Effective biometric technology used with big data," in Proceedings of Seventh International Congress on Information and Communication Technology: ICICT 2022, London, Volume 3, 2022, pp. 239-250.

[30] K. A. Al-Karawi, "Real-time adaptive training for forensic speaker verification in reverberation conditions," International Journal of Speech Technology vol. 26, pp. 1079-1089, 2023.

[31] F. Al-Dhaher, D. Y. Mohammed, M. Khalaf, K. Al-Karawi, M. Sarfraz, and M. M. Al Maathidi, "Real-Time Lie-Speech Determination Using Voice-Stress Technology," Iraqi Journal For Computer Science and Mathematics, vol. 5, pp. 81-93, 2024.

[32] L. P. Wong and M. Russell, "Text-dependent speaker verification under noisy conditions using parallel model combination," in Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP'01). 2001 IEEE International Conference on, 2001, pp. 457-460.

[33] H. Semary, K. A. Al-Karawi, M. M. Abdelwahab, and A. Elshabrawy, "A Review on Internet of Things (IoT)-Related Disabilities and Their Implications," Journal of Disability Research, vol. 3, pp. 1–16, 2024.

[34] K. A. Al-Karawi and A. S. Alenizi, "Reverberation Time and Distance Impact on the Equal Error Rate," in Proceedings of Ninth International Congress on Information and Communication Technology: ICICT 2024, London, Volume 10, 2024, p. 13.

[35] A. S. Alenizi and K. A. Al-Karawi, "Speaker Recognition with Deep Learning Approaches: A Review," in International Congress on Information and Communication Technology, 2024, pp. 481-499.

[36] H. Semary, K. A. Al-Karawi, and M. M. Abdelwahab, "Using voice technologies to support disabled people," Journal of Disability Research, vol. 3, pp. 1-8, 2024.

[37] K. A. Al-karawi, "Real-time adaptive training for forensic speaker verification in reverberation conditions," International Journal of Speech Technology, pp. 1-11, 2023.

[38] M. Hamidi, H. Satori, O. Zealouk, K. Satori, and N. Laaidi, "Interactive voice response server voice network administration using hidden markov model speech recognition system," in 2018 Second World Conference on Smart Trends in Systems, Security and Sustainability (WorldS4), 2018, pp. 16-21.

[39] R. Singh, B. Raj, and R. M. Stern, "Model compensation and matched condition methods for robust speech recognition," in Noise reduction in speech applications, ed: CRC press, 2018, pp. 245-275.

[40] L. Docio-Fernandez and C. Garcia-Mateo, "Noise model selection for robust speech recognition," in Fifth International Conference on Spoken Language Processing, 1998.

[41] K. Al-Karawi, "Robust speaker recognition in reverberant condition-toward greater biometric security," University of Salford, 2018.

Downloads

Published

2026-07-15

Issue

Section

Articles