A Decision Support System for Major Depressive Disorder Detection Using EMD and CNN-BiLSTM-Based Analysis of Human Voice.

Major depressive disorder (MDD) is a serious mental health disorder that is understood to affect an individual's speech signals, including observed variations in acoustic, prosodic, and laryngeal characteristics of a person's speech. There is a need for noninvasive, objective, and automated diagnostic methods for the diagnosis of depression. In this study, after applying the Empirical Mode Decomposition (EMD) to speech signals, it is proposed to classify the extracted features using a hybrid deep learning model (CNN-BiLSTM) for the early diagnosis of MDD. Nonstationary speech signals were decomposed into 7 intrinsic mode functions (IMF) using the EMD method. A total of 325 acoustic features (MFCC, LPCC, formant frequencies, jitter, shimmer, spectral entropy, glottal parameters, etc) were extracted from both the raw speech signal and the 7 IMF segments and classified using the CNN-BiLSTM model. Performance of the proposed method was compared with the performances of five different machine learning models for the MODMA dataset: Fine Tree, Linear SVM, Ensemble Boosted Trees, Kernel Logistic Regression, and Narrow Neural Network. The results show that EMD-based adaptive features significantly improve classification performance compared with features extracted without EMD. The study demonstrated better performance than MFCC or spectrogram-based studies in the literature, with 94.7% accuracy and an average F1-score value of 0.95. The results of 20% hold-out validation and 5-fold cross-validation demonstrate that the proposed method offers high accuracy, stability, interpretability, and generalizability. Additionally, a MATLAB-based clinical decision support interface supporting the CNN-BiLSTM model for MDD diagnosis from speech signals has been developed.
Mental Health
Care/Management

Authors

Gulenc Gulenc, Ozturk Ozturk
View on Pubmed
Share
Facebook
X (Twitter)
Bluesky
Linkedin
Copy to clipboard