Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
266
result(s) for
"Schuller, Björn W."
Sort by:
Representation learning with parameterised quantum circuits for advancing speech emotion recognition
2025
Quantum machine learning (QML) offers a promising avenue for advancing representation learning in complex signal domains. In this study, we investigate the use of parameterised quantum circuits (PQCs) for speech emotion recognition (SER)–a challenging task due to the subtle temporal variations and overlapping affective states in vocal signals. We propose a hybrid quantum-classical architecture that integrates PQCs into a conventional convolutional neural network (CNN), leveraging quantum properties such as superposition and entanglement to enrich emotional feature representations. Experimental evaluations on three benchmark datasets IEMOCAP, RECOLA, and MSP-IMPROV–demonstrate that our hybrid model achieves improved classification performance relative to a purely classical CNN baseline, with over 50% reduction in trainable parameters. Furthermore, Adjusted Rand Index (ARI) analysis demonstrates that the quantum model yields feature representations with improved alignment to true emotion classes compared with the classical model, reinforcing the observed performance gains. This work provides early evidence of the potential for QML to enhance emotion recognition and lays the foundation for future quantum-enabled affective computing systems.
Journal Article
A frequency analysis of filterbank initialisation and noise augmentation for LEAF
by
Rampp, Simon D. N.
,
Milling, Manuel
,
Schuller, Björn W.
in
631/114/1305
,
631/114/2397
,
639/166/987
2026
Differentiable frontends, such as the LEArnable Frontend (LEAF), have drawn increasing interest from the computer audition (CA) community combining the rigour of traditional signal processing techniques with the flexibility and potential of end-to-end deep learning approaches. Concretely, they promise the ability to automatically learn task-specific features, resulting in both higher performance and better interpretability of CA applications. With the adaptability of LEAF’s parameters being questioned in recent literature, we further dig into the reasons why LEAF does not adjust its parameters. We thus perform a detailed analysis investigating the effects of filterbank initialisation for LEAF in a wide, previously unmatched range of computer audition tasks, namely speech recognition, speech emotion recognition, acoustic scene classification, and bird activity detection. In line with literature, we report that performance stays constantly high irrespective of filterbank initialisation, so long as it covers the entire frequency spectrum, in which case adaptation is minimal. Crucially, however, a filterbank initialised with all frequency bands equally does change its centre frequencies and bandwidths, yet remains with a lower performance. This effect is seemingly independent of how information is spread across frequencies, as we confirm in an additional set of experiments with controlled frequency distributions. This points towards the critical role of initialisation and the inductive bias of LEAF and manifests concerns about the adaptability and interpretability of LEAF across many settings. The code for our experiments is publicly available under
https://github.com/millinma/LEAFFrequencyAnalysis
.
Journal Article
Explainable detection of machine generated music and early systematic evaluation
2026
Machine-generated music (MGM) has become a groundbreaking innovation with wide-ranging applications, such as music therapy, personalised editing, and creative inspiration within the music industry. However, the unregulated proliferation of MGM presents considerable challenges to the entertainment, education, and arts sectors by potentially undermining the value of high-quality human compositions. Consequently, MGM detection (MGMD) is crucial for preserving the integrity of these fields. Despite its significance, MGMD domain lacks comprehensive systematic evaluation results necessary to drive meaningful progress. To address this gap, we conduct experiments on existing large-scale datasets using a range of foundational models for audio processing, establishing systematic evaluation results tailored to the MGMD task. Our selection includes traditional machine learning models, deep neural networks, Transformer-based architectures, and State space models (SSM). Recognising the inherently multimodal nature of music, which integrates both melody and lyrics, we also explore fundamental multimodal models in our experiments. Beyond providing basic binary classification outcomes, we delve deeper into model behaviour using multiple explainable Artificial Intelligence (XAI) tools, offering insights into their decision-making processes. Our analysis reveals that ResNet18 performs the best according to in-domain and out-of-domain tests. By providing a comprehensive comparison of systematic evaluation results and their interpretability, we propose several directions to inspire future research to develop more robust and effective detection methods for MGM. We provide our codes and some samples on Github repository
https://github.com/myxp-lyp/Detecting-Machine-Generated-Music-with-Explainability-A-Challenge-and-Systematic-Evaluation
.
Journal Article
Prediction on Mechanical Properties of Non-Equiatomic High-Entropy Alloy by Atomistic Simulation and Machine Learning
by
Shibuta, Yasushi
,
Zhang, Liang
,
Qian, Kun
in
Alloys
,
Artificial neural networks
,
Cognitive tasks
2021
High-entropy alloys (HEAs) with multiple constituent elements have been extensively studied in the past 20 years, due to their promising engineering application. Previous experimental and computational studies of HEAs focused mainly on equiatomic or near equiatomic HEAs. However, there is probably far more treasure in those non-equiatomic HEAs with carefully designed composition. In this study, the molecular dynamics (MD) simulation combined with machine learning (ML) methods was used to predict the mechanical properties of non-equiatomic CuFeNiCrCo HEAs. A database was established based on a tensile test of 900 HEA single-crystal samples by MD simulation. Eight ML models were investigated and compared for the binary classification learning tasks, ranging from shallow models to deep models. It was found that the kernel-based extreme learning machine (KELM) model outperformed others for the prediction of yield stress and Young’s modulus. The accuracy of the KELM model was further verified by the large-sized polycrystal HEA samples. The results show that computational simulation combined with ML methods is an efficient way to predict the mechanical performance of HEAs, which provides new ideas for accelerating the development of novel alloy materials for engineering applications.
Journal Article
Facial Emotion Recognition of 16 Distinct Emotions From Smartphone Videos: Comparative Study of Machine Learning and Human Performance
2025
The development of automatic emotion recognition models from smartphone videos is a crucial step toward the dissemination of psychotherapeutic app interventions that encourage emotional expressions. Existing models focus mainly on the 6 basic emotions while neglecting other therapeutically relevant emotions. To support this research, we introduce the novel Stress Reduction Training Through the Recognition of Emotions Wizard-of-Oz (STREs WoZ) dataset, which contains facial videos of 16 distinct, therapeutically relevant emotions.
This study aimed to develop deep learning-based automatic facial emotion recognition (FER) models for binary (positive vs negative) and multiclass emotion classification tasks, assess the models' performance, and validate them by comparing the models with human observers.
The STREs WoZ dataset contains 14,412 facial videos of 63 individuals displaying the 16 emotions. The selfie-style videos were recorded during a stress reduction training using front-facing smartphone cameras in a nonconstrained laboratory setting. Automatic FER models using both appearance and deep-learned features for binary and multiclass emotion classification were trained on the STREs WoZ dataset. The appearance features were based on the Facial Action Coding System and extracted with OpenFace. The deep-learned features were obtained through a ResNet50 model. For our deep learning models, we used the appearance features, the deep-learned features, and their concatenation as inputs. We used 3 recurrent neural network (RNN)-based architectures: RNN-convolution, RNN-attention, and RNN-average networks. For validation, 3 human observers were also trained in binary and multiclass emotion recognition. A test set of 3018 facial emotion videos of the 16 emotions was completed by both the automatic FER model and human observers. The performance was assessed with unweighted average recall (UAR) and accuracy.
Models using appearance features outperformed those using deep-learned features, as well as models combining both feature types in both tasks, with the attention network using appearance features emerging as the best-performing model. The attention network achieved a UAR of 92.9% in the binary classification task, and accuracy values ranged from 59.0% to 90.0% in the multiclass classification task. Human performance was comparable to that of the automatic FER model in the binary classification task, with a UAR of 91.0%, and superior in the multiclass classification task, with accuracy values ranging from 87.4% to 99.8%.
Future studies are needed to enhance the performance of automatic FER models for practical use in psychotherapeutic apps. Nevertheless, this study represents an important first step toward advancing emotion-focused psychotherapeutic interventions via smartphone apps.
Journal Article
Voice Analysis for Neurological Disorder Recognition–A Systematic Review and Perspective on Emerging Trends
2022
Quantifying neurological disorders from voice is a rapidly growing field of research and holds promise for unobtrusive and large-scale disorder monitoring. The data recording setup and data analysis pipelines are both crucial aspects to effectively obtain relevant information from participants. Therefore, we performed a systematic review to provide a high-level overview of practices across various neurological disorders and highlight emerging trends. PRISMA-based literature searches were conducted through PubMed, Web of Science, and IEEE Xplore to identify publications in which original (i.e., newly recorded) datasets were collected. Disorders of interest were psychiatric as well as neurodegenerative disorders, such as bipolar disorder, depression, and stress, as well as amyotrophic lateral sclerosis amyotrophic lateral sclerosis, Alzheimer's, and Parkinson's disease, and speech impairments (aphasia, dysarthria, and dysphonia). Of the 43 retrieved studies, Parkinson's disease is represented most prominently with 19 discovered datasets. Free speech and read speech tasks are most commonly used across disorders. Besides popular feature extraction toolkits, many studies utilise custom-built feature sets. Correlations of acoustic features with psychiatric and neurodegenerative disorders are presented. In terms of analysis, statistical analysis for significance of individual features is commonly used, as well as predictive modeling approaches, especially with support vector machines and a small number of artificial neural networks. An emerging trend and recommendation for future studies is to collect data in everyday life to facilitate longitudinal data collection and to capture the behavior of participants more naturally. Another emerging trend is to record additional modalities to voice, which can potentially increase analytical performance.
Journal Article
An ecological momentary assessment study assessing repetitive negative thinking as a predictor for psychopathology
2025
Repetitive negative thinking (RNT), an important transdiagnostic process, is commonly assessed using trait questionnaires. While these instruments ask respondents to estimate their general tendency towards RNT, ecological momentary assessment (EMA) allows to assess how much individuals actually engage in RNT in their daily lives. In a sample of N = 1,176 adolescents and young adults, we investigated whether average levels of RNT assessed via EMA predicted psychopathological symptoms. Adjusting for trait RNT measures and baseline scores on outcome measures, we found that average levels of RNT assessed via EMA significantly predicted higher depressive and anxiety symptoms as well as lower mental well-being at baseline, one-, three-, and twelve-month follow-up. Exploratory analyses of the association between temporal dynamics of RNT (e.g., RNT inertia) and psychopathological symptoms yielded inconsistent results. The high predictive power of average scores on the EMA-based RNT measure suggests that EMA is a promising tool for assessing RNT.
Journal Article
Emotion recognition in live broadcasting: a multimodal deep learning framework
by
Li, Xuewei
,
Schuller, Björn W
,
Abbas, Rizwan
in
Artificial neural networks
,
Broadcasting
,
Computational efficiency
2025
Multimodal emotion recognition is a rapidly developing field with applications across diverse fields such as entertainment, healthcare, marketing, and education. The emergence of live broadcasting demands real-time emotion recognition, which involves analyzing emotions via body language, voice, facial expressions, and context. Previous studies have faced challenges associated with multimodal emotion recognition in live broadcasting, such as computational efficiency, noisy and incomplete data, and difficult camera angles. This research presents a Multimodal Emotion Recognition in Live Broadcasting (MERLB) system that collects speech, facial expressions, and context displayed in live broadcasting for emotion recognition. We utilize a deep convolutional neural network architecture for facial emotion recognition, incorporating inception modules and dense blocks. We aim to enhance computational efficiency by focusing on key segments rather than analyzing the entire utterance. MERLB employs tensor train layers to combine multimodal representations at higher orders. Experiments were conducted on the FIFA, League of Legends, IEMOCAP, and CMU-MOSEI datasets. MERLB achieves a 6.44% F1 score improvement on the FIFA dataset and 4.71% on League of Legends, outperforming other multi-modal emotion methods on IEMOCAP and CMU-MOSEI datasets. Our code is available at https://github.com/swerizwan/merlb.
Journal Article
CovNet: A Transfer Learning Framework for Automatic COVID-19 Detection From Crowd-Sourced Cough Sounds
Since the COronaVIrus Disease 2019 (COVID-19) outbreak, developing a digital diagnostic tool to detect COVID-19 from respiratory sounds with computer audition has become an essential topic due to its advantages of being swift, low-cost, and eco-friendly. However, prior studies mainly focused on small-scale COVID-19 datasets. To build a robust model, the large-scale multi-sound FluSense dataset is utilised to help detect COVID-19 from cough sounds in this study. Due to the gap between FluSense and the COVID-19-related datasets consisting of cough only, the transfer learning framework (namely CovNet) is proposed and applied rather than simply augmenting the training data with FluSense. The CovNet contains (i) a parameter transferring strategy and (ii) an embedding incorporation strategy. Specifically, to validate the CovNet's effectiveness, it is used to transfer knowledge from FluSense to COUGHVID, a large-scale cough sound database of COVID-19 negative and COVID-19 positive individuals. The trained model on FluSense and COUGHVID is further applied under the CovNet to another two small-scale cough datasets for COVID-19 detection, the COVID-19 cough sub-challenge (CCS) database in the INTERSPEECH Computational Paralinguistics challengE (ComParE) challenge and the DiCOVA Track-1 database. By training four simple convolutional neural networks (CNNs) in the transfer learning framework, our approach achieves an absolute improvement of 3.57% over the baseline of DiCOVA Track-1 validation of the area under the receiver operating characteristic curve (ROC AUC) and an absolute improvement of 1.73% over the baseline of ComParE CCS test unweighted average recall (UAR).
Journal Article