Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
75
result(s) for
"Stanley, Emma A M"
Sort by:
Connecting algorithmic fairness and fair outcomes in a sociotechnical simulation case study of AI-assisted healthcare
by
Wilms, Matthias
,
Gillett, Haley
,
Vigneshwaran, Vibujithan
in
639/166/985
,
639/705/1041
,
692/700/3935
2025
Artificial intelligence (AI) has vast potential for improving healthcare delivery, but concerns regarding biases in these systems have raised important questions regarding fairness when deployed clinically. Most prior studies on fairness in clinical AI focus solely on performance disparities between subpopulations, which often fall short of connecting the technical outputs of AI systems with sociotechnical outcomes. In this work, we present a simulation-based approach to explore how statistical definitions of algorithmic fairness translate to fairness in long-term outcomes, using AI-assisted breast cancer screening as a case example. We evaluate four fairness criteria and their impact on mortality rates and socioeconomic disparities, while also considering how clinical decision makers’ reliance on AI and patients’ access to healthcare affect outcomes. Our results highlight how algorithmic fairness does not directly translate into fair and equitable outcomes, underscoring the importance of integrating sociotechnical perspectives to gain a holistic understanding of fairness in healthcare AI.
Artificial intelligence (AI) can greatly improve healthcare delivery and outcomes, but potential embedded biases can affect fairness in clinical deployment. Here, the authors develop a simulation-based approach to explore which formalisations of AI algorithmic fairness translate into long-term outcome fairness, with a focus on breast cancer.
Journal Article
Distinct visual biases affect humans and artificial intelligence in medical imaging diagnoses
by
McLeod, Graham A.
,
Forkert, Nils D.
,
Stanley, Emma A. M.
in
631/114
,
631/114/1314
,
631/378/2613/2615
2025
Artificial intelligence (AI) systems can detect subtle features in diagnostic imaging scans that radiologists may miss, including higher-order features that lack obvious visual correlates. This may enable earlier disease detection and non-invasive lesion phenotyping, but also introduces risks due to AI’s reliance on correlations rather than causation, potential demographic and technical biases, and uninterpretable reasoning. This perspective explores how radiologists and AI learn to perceive details in medical images differently, leading to potential discrepancies in medical decision-making.
Journal Article
Combining federated learning and travelling model boosts performance and opens opportunities for digital health equity
2026
Federated learning (FL) and travelling model (TM) allow privacy-preserving model training across sites without sharing patient-sensitive data. While both approaches have shown success, they face unique challenges related to distribution shifts between sites. To address this, we propose FedTM, a hybrid framework combining the strengths of FL and TM. FedTM begins with FL warmup training at sites with larger datasets, followed by sequential refinement through TM across all sites. We evaluated FedTM for Parkinson’s disease classification using 1817 brain scans from 83 international sites. Model performance, misclassification disparities, and communication costs were computed and compared to standard FL and TM approaches. Our results reveal that FedTM improves AUROC from 77 ± 0.01% to 82 ± 0.01%, reduces misclassification disparities from 34 ± 0.01% to 26 ± 0.01%, and decreases training load for smaller sites from 22 to 12 cycles. These advancements mark an important step toward promoting global healthcare equity and advancing responsible AI development.
Journal Article
Brain Aging in Patients With Cardiovascular Disease From the UK Biobank
2025
The brain undergoes complex but normal structural changes during the aging process in healthy adults, whereas deviations from the normal aging patterns of the brain can be indicative of various conditions as well as an increased risk for the development of diseases. The brain age gap (BAG), which is defined as the difference between the chronological age and the machine learning‐predicted biological age of an individual, is a promising biomarker for determining whether an individual deviates from normal brain aging patterns. While the BAG has shown promise for various neurological diseases and cardiovascular risk factors, its utility to quantify brain changes associated with diagnosed cardiovascular diseases has not been investigated to date, which is the aim of this study. T1‐weighted MRI scans from healthy participants in the UK Biobank were used to train a convolutional neural network (CNN) model for biological brain age prediction. The trained model was then used to quantify and compare the BAGs for all participants in the UK Biobank with known cardiovascular diseases, as well as healthy controls and patients with known neurological diseases for benchmark comparisons. Saliency maps were computed for each individual to investigate whether brain regions used for biological brain age prediction by the CNN differ between groups. The analyses revealed significant differences in BAG distributions for 10 of the 42 sex‐specific cardiovascular disease groups investigated compared to healthy participants, indicating disease‐specific variations in brain aging. However, no significant differences were found regarding the brain regions used for brain age prediction as determined by saliency maps, indicating that the model mostly relied on healthy brain aging patterns, even in the presence of cardiovascular diseases. Overall, the findings of this work demonstrate that the BAG is a sensitive imaging biomarker to detect differences in brain aging associated with specific cardiovascular diseases. This further supports the theory of the heart–brain axis by exemplifying that many cardiovascular diseases are associated with atypical brain aging. This study reveals that in patients with certain cardiovascular diseases there is atypical biological brain age predicted by an artificial intelligence model from T1‐weighted brain MRIs. This suggests the existence of distinct aging‐like patterns within the brain in the precense of cardiovascular conditions. These findings show the complex interaction between cardiovascular health and neurological health.
Journal Article
Self-supervised identification and elimination of harmful datasets in distributed machine learning for medical image analysis
by
Wilms, Matthias
,
Dagasso, Gabrielle
,
Ohara, Erik Y.
in
639/166/985
,
639/705/1042
,
Biomedicine
2025
Distributed learning enables collaborative machine learning model training without requiring cross-institutional data sharing, thereby addressing privacy concerns. However, local quality control variability can negatively impact model performance while systematic human visual inspection is time-consuming and may violate the goal of keeping data inaccessible outside acquisition centers. This work proposes a novel self-supervised method to identify and eliminate harmful data during distributed learning model training fully-automatically. Harmful data is defined as samples that, when included in training, increase misdiagnosis rates. The method was tested using neuroimaging data from 83 centers for Parkinson’s disease classification with simulated inclusion of a few harmful data samples. The proposed method reliably identified harmful images, with centers providing only harmful datasets being easier to identify than single harmful images within otherwise good datasets. While only evaluated using neuroimaging data, the presented method is application-agnostic and presents a step towards automated quality control in distributed learning.
Journal Article
Towards realistic simulation of disease progression in the visual cortex with CNNs
by
Wilms, Matthias
,
Vigneshwaran, Vibujithan
,
Stanley, Emma A. M.
in
631/378/116
,
631/378/116/1925
,
692/699/375/132
2025
Convolutional neural networks (CNNs) and mammalian visual systems share architectural and information processing similarities. We leverage these parallels to develop an in-silico CNN model simulating diseases affecting the visual system. This model aims to replicate neural complexities in an experimentally controlled environment. Therefore, we examine object recognition and internal representations of a CNN under neurodegeneration and neuroplasticity conditions simulated through synaptic weight decay and retraining. This approach can model neurodegeneration from events like tau accumulation, reflecting cognitive decline in diseases such as posterior cortical atrophy, a condition that can accompany Alzheimer’s disease and primarily affects the visual system. After each degeneration iteration, we retrain unaffected synapses to simulate ongoing neuroplasticity. Our results show that with significant synaptic decay and limited retraining, the model’s representational similarity decreases compared to a healthy model. Early CNN layers retain high similarity to the healthy model, while later layers are more prone to degradation. The results of this study reveal a progressive decline in object recognition proficiency, mirroring posterior cortical atrophy progression. In-silico modeling of neurodegenerative diseases can enhance our understanding of disease progression and aid in developing targeted rehabilitation and treatments.
Journal Article
Challenges and Potential of Artificial Intelligence in Neuroradiology
by
Stanley, Emma AM
,
Forkert, Nils D.
,
Winder, Anthony J.
in
Artificial intelligence
,
Brain research
,
Cost analysis
2024
Purpose
Artificial intelligence (AI) has emerged as a transformative force in medical research and is garnering increased attention in the public consciousness. This represents a critical time period in which medical researchers, healthcare providers, insurers, regulatory agencies, and patients are all developing and shaping their beliefs and policies regarding the use of AI in the healthcare sector. The successful deployment of AI will require support from all these groups. This commentary proposes that widespread support for medical AI must be driven by clear and transparent scientific reporting, beginning at the earliest stages of scientific research.
Methods
A review of relevant guidelines and literature describing how scientific reporting plays a central role at key stages in the life cycle of an AI software product was conducted. To contextualize this principle within a specific medical domain, we discuss the current state of predictive tissue outcome modeling in acute ischemic stroke and the unique challenges presented therein.
Results and Conclusion
Translating AI methods from the research to the clinical domain is complicated by challenges related to model design and validation studies, medical product regulations, and healthcare providers’ reservations regarding AI’s efficacy and affordability. However, each of these limitations is also an opportunity for high-impact research that will help to accelerate the clinical adoption of state-of-the-art medical AI. In all cases, establishing and adhering to appropriate reporting standards is an important responsibility that is shared by all of the parties involved in the life cycle of a prospective AI software product.
Journal Article
Towards objective and systematic evaluation of bias in artificial intelligence for medical imaging
2024
Artificial intelligence (AI) models trained using medical images for clinical tasks often exhibit bias in the form of disparities in performance between subgroups. Since not all sources of biases in real-world medical imaging data are easily identifiable, it is challenging to comprehensively assess how those biases are encoded in models, and how capable bias mitigation methods are at ameliorating performance disparities. In this article, we introduce a novel analysis framework for systematically and objectively investigating the impact of biases in medical images on AI models. We developed and tested this framework for conducting controlled in silico trials to assess bias in medical imaging AI using a tool for generating synthetic magnetic resonance images with known disease effects and sources of bias. The feasibility is showcased by using three counterfactual bias scenarios to measure the impact of simulated bias effects on a convolutional neural network (CNN) classifier and the efficacy of three bias mitigation strategies. The analysis revealed that the simulated biases resulted in expected subgroup performance disparities when the CNN was trained on the synthetic datasets. Moreover, reweighing was identified as the most successful bias mitigation strategy for this setup, and we demonstrated how explainable AI methods can aid in investigating the manifestation of bias in the model using this framework. Developing fair AI models is a considerable challenge given that many and often unknown sources of biases can be present in medical imaging datasets. In this work, we present a novel methodology to objectively study the impact of biases and mitigation strategies on deep learning pipelines, which can support the development of clinical AI that is robust and responsible.
Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
2026
Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-RADS severity, and cancer status using a unified frozen-backbone linear-probe protocol, training on 3 source datasets and evaluating on 12 task-compatible out-of-distribution (OOD) datasets after label harmonization. Mammography-specific vision-language models (Mammo-FM and MaMA) provide the strongest mean OOD performance, but robustness is not explained by mammography exposure alone. DINOv3 remains a competitive vision-only baseline, and mammography-adapted pretraining does not consistently improve generalization. Dataset-level analysis further shows that even leading models show heterogeneous performance across datasets. Feature-space inspection reveals that useful representations can preserve clinical signal while retaining dataset and acquisition structure. These findings highlight dataset-level OOD evaluation as a central criterion for assessing mammography representations. Our code is publicly available: https://github.com/biomedia-mira/mammo-ood.
A Neuroimaging Simulation Framework for Developing and Evaluating Causal AI
by
Wilms, Matthias
,
Vigneshwaran, Vibujithan
,
Ohara, Erik Y
in
Artificial intelligence
,
Biomarkers
,
Errors
2026
Causally linking disease-related factors to image-derived biomarkers provides a powerful pathway to understanding disease mechanisms. Despite growing interest in applying causal artificial intelligence (AI) approaches for this task, these methods still need to be adapted for complex medical images, and especially, neuroimaging. However, the lack of ground-truth data presents a barrier to development. To bridge this gap, we developed and tested a method for generating synthetic neuroimages, which adhere to a user-specified causal structure describing the non-image to image variable relationships, permitting the creation of ground-truth neuroimaging datasets. In the simulated T1-weighted magnetic resonance images, anatomical variability is modeled by sampling from a subspace estimated from real data and deforming a template image to create unique simulated subjects. Causal relationships are encoded via precise volumetric changes of any region-of-interest without unwanted global artifacts. We achieved relative volume errors of 0.3-2.66% for the targeted regions-of-interest and demonstrate their statistically significant causal relationships, while maintaining mean absolute errors for non-target brain regions between 0.034-0.397ml. An initial evaluation of causal discovery methods exposes their limited ability to suppress spurious connections, highlighting the need for image-appropriate methods. Our framework is the first to enable the generation of realistic synthetic 3D neuroimages with explicit causal control that can serve as the missing ground-truth data necessary for the objective benchmarking and development of causal AI methods.