Abstract:
The advancement of Deep Learning has revolutionised numerous fields, including humancomputer
interaction (HCI) and emotion detection. In speech recognition, voice assistants such
as Apple's Siri, Microsoft's Cortana, and Google's Assistant have become integral to everyday
technology use. However, current emotion detection systems often need more support for
African languages, presenting a significant research gap. The aim and objectives of the study
are to develop a Speech Emotion Recognition (SER) system tailored to South African
languages, with a specific focus on Afrikaans. The primary objective is to compile an Afrikaans
speech corpus for emotion recognition. Furthermore, it explores hybrid neural network
architectures combining CNN, Recurrent Neural Networks (RNNs) and Long Short-Term
Memory (LSTM) networks for improved accuracy. Finally, the aim is to investigate the
optimisation of SER models for multilingual cross-language support.
This study adopts a pragmatic research paradigm, utilising the Design Science Research
Framework (DSR). Data is collected quantitatively. The literature review covers human
speech, speech physiology, neural network architectures, speech corpus, and data extraction
methods. Hybrid architectures and preprocessing techniques are examined to identify the most
effective configurations. The Afrikaans speech corpus is sourced from Creative Commons
(CC) sources, with data augmentation using Generative Adversarial Networks (GANs).
The study successfully compiled an Afrikaans speech corpus. Synthetic speech samples were
created using frameworks such as Tacotron 2 and the WaveNet vocoder, thereby enhancing
sound quality. For the Hybrid Neural Network Architectures: The novel Dendritic
Convolutional Long Short-Term Memory (DCLSTM) architecture outperformed traditional
CNN and LSTM models, effectively capturing nuanced and long-term emotional dependencies
in audio data. The second Kalman filter variant was identified as the optimal preprocessing
technique for noisy speech signals. A final accuracy of 78.50% was achieved, improving on
baseline CNN and LSTM models. For the Multilingual SER System, the DenCaps model
achieved up to 71.91% accuracy on multilingual datasets, demonstrating improved
generalisation and better capture of temporal dynamics and spatial hierarchies in emotional
speech.
The findings highlight the effectiveness of hybrid neural network architectures and advanced
preprocessing techniques in improving SER accuracy. The development of the Afrikaans
speech corpus and the introduction of the DenCaps model represent significant advancements
in emotion detection. These results underscore the potential for creating robust, multilingual
SER systems that cater to diverse linguistic needs. The research objectives were successfully
achieved through the development of a specialised Afrikaans speech corpus, the exploration of
innovative neural network architectures, and the optimisation of models for multilingual
support. The academic contributions pave the way for more accurate, contextually aware
human-computer interaction systems that address the research gap in emotion detection for African languages.