Noise-Robust Speech Emotion Recognition using MFCC and CNN with Temporal Emotional Analysis
ID:18
View protection:Participant Only
Updated time:2026-07-22 16:09:05 Views:11
Online
Abstract
The SER system holds an important position in simplifying the process of human-computer interaction by recognizing human emotions. In this research paper, we introduce an innovative approach for developing an efficient speech emotion recognition system, which involves the use of preprocessing methods and deep learning algorithms for classifying speech into emotional classes without any interference from noises and nonspeech data. The proposed algorithm uses voice activity detection (VAD) to filter out speech data only and perform noise reduction to minimize interference while keeping intact emotion-related data in the filtered voice segment. Next, we extract several features of speech, which include Mel-frequency cepstral coefficients (MFCCs) and other spectral features like chroma, melspectrogram, and spectral contrast to capture different aspects of speech. We implement deep learning models with convolutional networks for emotion classification. Finally, the system performs time-series emotional analysis by breaking down audio signal samples into one-second samples for tracking any changes in emotions over time. The visual representation methods of the proposed approach include waveform visualization, spectrogram and emotional timeline.
Keywords
SER,MFCC,CNN,Noise Reduction,Feature Extraction,Spectrogram Analysis,Temporal Analysis
Post comments