Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
6
result(s) for
"Warule, Pankaj"
Sort by:
Speech emotion recognition using MFCC-based entropy feature
2024
The prime objective of speech emotion recognition is to accurately recognize the emotion from the speech signal. It is a challenging task to accomplish. Speech emotion recognition (SER) has many applications, including medicine, online marketing, strengthening human–computer interaction (HCI), online education, and many more. Hence, it has been a topic of interest for many researchers for last three decades. The researchers used different methodologies to improve the classification accuracy of emotions. In this study, we tried to improve emotion classification accuracy using mel-frequency cepstral coefficient (MFCC)-based entropy features. First, we extracted the MFCC coefficient matrix from every speech of the EMO-DB, RAVDESS and SAVEE datasets, and then we calculated the proposed features: statistical mean (
MFCC
mean
), MFCC-based approximate entropy (
MFCC
AE
), and MFCC-based spectral entropy (
MFCC
SE
), from the MFCC coefficient matrix of every utterance. The performance of the proposed features is accessed using the DNN classifier. We achieved a classification accuracy of 87.48%, 75.9%, and 79.64% using the combination of
MFCC
mean
and
MFCC
SE
features and obtained classification accuracies of 85.61%, 77.54%, and 76.26% using the combination of
MFCC
mean
,
MFCC
AE
, and
MFCC
SE
features for the EMO-DB, RAVDESS, and SAVEE datasets, respectively.
Journal Article
Time-frequency analysis of speech signal using Chirplet transform for automatic diagnosis of Parkinson’s disease
by
Mishra, Siba Prasad
,
Deb, Suman
,
Warule, Pankaj
in
Biological and Medical Physics
,
Biomedical Engineering and Bioengineering
,
Biomedicine
2023
Parkinson’s disease (PD) is the second most prevalent neurodegenerative disorder in the world after Alzheimer’s disease. Early diagnosing PD is challenging as it evolved slowly, and its symptoms eventuate gradually. Recent studies have demonstrated that changes in speech may be utilized as an excellent biomarker for the early diagnosis of PD. In this study, we have proposed a Chirplet transform (CT) based novel approach for diagnosing PD using speech signals. We employed CT to get the time-frequency matrix (TFM) of each speech recording, and we extracted time-frequency based entropy (TFE) features from the TFM. The statistical analysis demonstrates that the TFE features reflect the changes in speech that occurs in the speech due to PD, hence can be used for classifying the PD and healthy control (HC) individuals. The effectiveness of the proposed framework is validated using the vowels and words from the PC-GITA database. The genetic algorithm is utilized to select the optimum features subset, while a support vector machine (SVM), decision tree (DT), K-Nearest Neighbor (KNN), and Naïve Bayes (NB) classifiers are employed for classification. The TFE features outperform the breathiness and Mel frequency cepstral coefficients (MFCC) features. The SVM classifier is most effective compared to other machine-learning classifiers. The highest classification accuracy rates of 98% and 99% are achieved using the vowel /a/ and word /atleta/, respectively. The results reveal that the proposed CT-based entropy features effectively diagnose PD using the speech of a person.
Journal Article
Speech emotion classification using feature-level and classifier-level fusion
2024
Emotion plays a vital role in every living being. Understanding emotion is a very complex task for everyone, but if possible, it will work like a miracle to solve thousands of problems and save many lives. Emotion is reflected not only in the gesture but also in work and in producing an efficient result. Hence, the recognition of emotion using speech has been a topic of interest for many researchers for the last three decades. In our study, we used three features, mel frequency cepstral coefficient (MFCC), spectrogram, and mel-spectrogram, as a one-dimensional input vector to the convolutional neural network (CNN) and deep neural network (DNN) for speech emotion classification. We evaluated the accuracy of SER using the features individually and in combination with the deep learning classifiers CNN and DNN. For both CNN and DNN classifiers, the combination of features performed better than the individual features. The combination of features using the DNN classifier achieved an accuracy of 76.60%, 87.10%, 79.79%, and 100%, and using the CNN classifier achieved classification accuracy of 75%, 84.11%, 78.13%, and 100% for the RAVDESS, EMO-DB, SAVEE, and TESS datasets,respectively. Then we applied a proposed feature and classifier-level fusion method using CNN and DNN to improve emotion classification performance and achieved classification accuracy of 80.42%, 87.48%, and 80.99% on the RAVDESS, EMO-DB, and SAVEE datasets, respectively. The performance of the proposed feature and classifier-level fusion method was compared with the other methods, and it was found that the proposed method performed better than the state-of-the-art methods.
Graphical abstract
Journal Article
Detection of Common Cold from Speech Signals using Deep Neural Network
2023
This paper presents a deep learning-based analysis and classification of cold speech observed when a person is diagnosed with the common cold. The common cold is a viral infectious disease that affects the throat and the nose. Since speech is produced by the vocal tract after linear filtering of excitation source information, during a common cold, its attributes are impacted by the throat and the nose. The proposed study attempts to develop a deep learning-based classification model that can accurately predict whether a person has a cold or not based on their speech. The common cold-related information is captured using Mel-frequency cepstral coefficients (MFCC) and linear predictive coding (LPC) from the speech signal. The data imbalance is handled using the sampling strategy, SMOTE–Tomek links. Then, utilizing MFCC and LPC features, a deep learning-based model is trained and then used to categorize cold speech. The performance of a deep learning-based method is compared to logistic regression, random forest, and gradient boosted tree classifiers. The proposed model is less complex and uses a smaller feature set while giving comparable results to other state-of-the-art methods. The proposed method gives an UAR of 67.71%, higher than the benchmark OpenSMILE SVM result of 64%. The study’s success will yield a noninvasive method for cold detection, which can further be extended to detect other speech-affecting pathologies.
Journal Article
Significance of voiced and unvoiced speech segments for the detection of common cold
by
Mishra, Siba Prasad
,
Deb, Suman
,
Warule, Pankaj
in
Classification
,
Computer Imaging
,
Computer Science
2023
This work investigates the significance of the voiced and unvoiced region for detecting common cold from the speech signal. In literature, the entire speech signal is processed to detect the common cold and other diseases. This study uses a short-time energy-based approach to segment the voiced and unvoiced region of the speech signal. Then, frame-wise mel frequency cepstral coefficients (MFCC) features are extracted from the voiced and unvoiced segments of each speech utterance, and statistics (mean, variance, skewness, and kurtosis) are calculated to get the feature vector for each speech utterance. The support vector machine (SVM) is utilized to analyze the performance of features extracted from the voiced and unvoiced region. Result shows that the feature extracted from voiced segments, unvoiced segments, and complete active speech (CAS) gives almost similar results using the MFCC features and SVM classifier. Therefore, rather than processing the CAS, we can process the unvoiced speech segments, which have fewer frames compared to CAS and voiced regions of speech. The processing of solely unvoiced segments can reduce the time and computation complexity of a speech signal-based common cold detection system.
Journal Article
Dual-Tree Complex Wavelet Transform for the Automatic Detection of the Common Cold Based on Speech Signals
by
Mishra, Siba Prasad
,
Deb, Suman
,
Warule, Pankaj
in
Biomedical engineering
,
Biomedicine
,
Circuits and Systems
2025
The acoustic and prosodic features of speech change in the presence of various health states. Biomedical engineering has enormous promise for developing non-invasive diagnostic technologies that use voice as a modality. The common cold is a highly prevalent sickness that affects a significant proportion of the global population throughout the year. The utilization of speech signals for the detection of the common cold has experienced a surge in popularity in recent times. In this study, the dual-tree complex wavelet transform (DTCWT) based new feature extraction technique is proposed for diagnosing common cold infection. First, we have employed the DTCWT to break down the speech signal into many sub-band coefficients. Then the features such as mean, variance, skewness, kurtosis, energy, approximate entropy, Renyi entropy, and permutation entropy are extracted from these sub-band coefficients. The URTIC database is utilized to assess the effectiveness of the proposed features. The classification results achieved using the transformer model show that the proposed algorithm detected cold from a speech sample with UAR of 68.66% and 64.52% on the develop and test set of the URTIC dataset. We have obtained comparable results with the state-of-the-art methods. The DTCWT captures subtle changes in speech signals, making it well-suited for detecting common cold symptoms. Its ability to provide both time-frequency localization and phase information enables it to discriminate between healthy and cold-affected speech patterns, leading to improved classification accuracy.
Journal Article