Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Series TitleSeries Title
-
Reading LevelReading Level
-
YearFrom:-To:
-
More FiltersMore FiltersContent TypeItem TypeIs Full-Text AvailableSubjectPublisherSourceDonorLanguagePlace of PublicationContributorsLocation
Done
Filters
Reset
479
result(s) for
"Robinson, Peter N"
Sort by:
Classification, Ontology, and Precision Medicine
by
Haendel, Melissa A
,
Robinson, Peter N
,
Chute, Christopher G
in
Classification
,
Data processing
,
Datasets as Topic
2018
Data-organizing methods have been in place for centuries, but very large data sets have come into being relatively recently. The authors describe terminologies, ontologies, and the changes needed to permit analyses of “big data” that might better serve medical decision making.
Journal Article
Biometric and structural ocular manifestations of Marfan syndrome
2017
To study biometric and structural ocular manifestations of Marfan syndrome (MFS).
Observational, retrospective, comparative cohort study in a tertiary referral center on 285 MFS patients and 267 controls. Structural and biometric ocular characteristic were compared.
MFS eyes were longer (axial length 24.25 ± 1.74 mm versus 23.89 ± 1.31 mm, p < 0.001) and had a flatter cornea than control eyes (mean keratometry 41.78 ± 1.80 diopters (D) versus 43.05 ± 1.51 D, p < 0.001). Corneal astigmatism was greater and the central cornea was thinner in MFS eyes (530.14 ± 41.31 μm versus 547.02 ± 39.18 μm, p < 0.001). MFS eyes were more myopic than control eyes (spherical equivalent -2.16 ± 3.75 D versus -1.17 ± 2.58 D, p < 0.001). Visual acuity was reduced (0.13 ± 0.25 logMAR versus 0.05 ± 0.18 logMAR, p < 0.001) and intraocular pressure was lower in MFS eyes (14.6 ± 3.4 mmHg versus 15.1 ± 3.2 mmHg, p = 0.01). Iris transillumination defects (ITD) were significantly more common in MFS eyes (odds ratio for MFS in the presence of ITD, 3.7). Ectopia lentis (EL) was only present in MFS eyes (33.4%). History of retinal detachment was significantly more common in MFS eyes. Glaucoma was equally common in both groups.
ITD and EL are most characteristic findings in MFS. ITD and corneal curvature should be studied as diagnostic criteria for MFS. Visual acuity is reduced in MFS. MFS patients need regular eye exams to identify serious ocular complications.
Journal Article
Phenotype-driven strategies for exome prioritization of human Mendelian disease genes
by
Robinson, Peter N.
,
Smedley, Damian
in
Bioinformatics
,
Biomedical and Life Sciences
,
Biomedicine
2015
Whole exome sequencing has altered the way in which rare diseases are diagnosed and disease genes identified. Hundreds of novel disease-associated genes have been characterized by whole exome sequencing in the past five years, yet the identification of disease-causing mutations is often challenging because of the large number of rare variants that are being revealed. Gene prioritization aims to rank the most probable candidate genes towards the top of a list of potentially pathogenic variants. A promising new approach involves the computational comparison of the phenotypic abnormalities of the individual being investigated with those previously associated with human diseases or genetically modified model organisms. In this review, we compare and contrast the strengths and weaknesses of current phenotype-driven computational algorithms, including Phevor, Phen-Gen, eXtasy and two algorithms developed by our groups called PhenIX and Exomiser. Computational phenotype analysis can substantially improve the performance of exome analysis pipelines.
Journal Article
Imbalance-Aware Machine Learning for Predicting Rare and Common Disease-Associated Non-Coding Variants
by
Robinson, Peter N.
,
Valentini, Giorgio
,
Re, Matteo
in
631/114/1305
,
631/114/2413
,
631/114/2785
2017
Disease and trait-associated variants represent a tiny minority of all known genetic variation, and therefore there is necessarily an imbalance between the small set of available disease-associated and the much larger set of non-deleterious genomic variation, especially in non-coding regulatory regions of human genome. Machine Learning (ML) methods for predicting disease-associated non-coding variants are faced with a chicken and egg problem - such variants cannot be easily found without ML, but ML cannot begin to be effective until a sufficient number of instances have been found. Most of state-of-the-art ML-based methods do not adopt specific imbalance-aware learning techniques to deal with imbalanced data that naturally arise in several genome-wide variant scoring problems, thus resulting in a significant reduction of sensitivity and precision. We present a novel method that adopts imbalance-aware learning strategies based on resampling techniques and a hyper-ensemble approach that outperforms state-of-the-art methods in two different contexts: the prediction of non-coding variants associated with Mendelian and with complex diseases. We show that imbalance-aware ML is a key issue for the design of robust and accurate prediction algorithms and we provide a method and an easy-to-use software tool that can be effectively applied to this challenging prediction task.
Journal Article
Phenotype Ontologies and Cross-Species Analysis for Translational Research
2014
The use of model organisms as tools for the investigation of human genetic variation has significantly and rapidly advanced our understanding of the aetiologies underlying hereditary traits. However, while equivalences in the DNA sequence of two species may be readily inferred through evolutionary models, the identification of equivalence in the phenotypic consequences resulting from comparable genetic variation is far from straightforward, limiting the value of the modelling paradigm. In this review, we provide an overview of the emerging statistical and computational approaches to objectively identify phenotypic equivalence between human and model organisms with examples from the vertebrate models, mouse and zebrafish. Firstly, we discuss enrichment approaches, which deem the most frequent phenotype among the orthologues of a set of genes associated with a common human phenotype as the orthologous phenotype, or phenolog, in the model species. Secondly, we introduce and discuss computational reasoning approaches to identify phenotypic equivalences made possible through the development of intra- and interspecies ontologies. Finally, we consider the particular challenges involved in modelling neuropsychiatric disorders, which illustrate many of the remaining difficulties in developing comprehensive and unequivocal interspecies phenotype mappings.
Journal Article
An evaluation of GPT models for phenotype concept recognition
by
Groza, Tudor
,
Robinson, Peter N.
,
Baynam, Gareth
in
Annotations
,
Artificial intelligence
,
Automation
2024
Objective
Clinical deep phenotyping and phenotype annotation play a critical role in both the diagnosis of patients with rare disorders as well as in building computationally-tractable knowledge in the rare disorders field. These processes rely on using ontology concepts, often from the Human Phenotype Ontology, in conjunction with a phenotype concept recognition task (supported usually by machine learning methods) to curate patient profiles or existing scientific literature. With the significant shift in the use of large language models (LLMs) for most NLP tasks, we examine the performance of the latest Generative Pre-trained Transformer (GPT) models underpinning ChatGPT as a foundation for the tasks of clinical phenotyping and phenotype annotation.
Materials and methods
The experimental setup of the study included seven prompts of various levels of specificity, two GPT models (gpt-3.5-turbo and gpt-4.0) and two established gold standard corpora for phenotype recognition, one consisting of publication abstracts and the other clinical observations.
Results
The best run, using in-context learning, achieved 0.58 document-level F1 score on publication abstracts and 0.75 document-level F1 score on clinical observations, as well as a mention-level F1 score of 0.7, which surpasses the current best in class tool. Without in-context learning, however, performance is significantly below the existing approaches.
Conclusion
Our experiments show that gpt-4.0 surpasses the state of the art performance if the task is constrained to a subset of the target ontology where there is prior knowledge of the terms that are expected to be matched. While the results are promising, the non-deterministic nature of the outcomes, the high cost and the lack of concordance between different runs using the same prompt and input make the use of these LLMs challenging for this particular task.
Journal Article
Replacing non-biomedical concepts improves embedding of biomedical concepts
by
Niyonkuru, Enock
,
Robinson, Peter N.
,
Reese, Justin T.
in
Algorithms
,
Clusters
,
Computational linguistics
2025
Embeddings are semantically meaningful representations of words in a vector space, commonly used to enhance downstream machine learning applications. Traditional biomedical embedding techniques often replace all synonymous words representing biological or medical concepts with a unique token, ensuring consistent representation and improving embedding quality. However, the potential impact of replacing non-biomedical concept synonyms has received less attention. Embedding approaches often employ concept replacement to replace concepts that span multiple words, such as non-small-cell lung carcinoma, with a single concept identifier (e.g., D002289). Also, all synonyms of each concept are merged into the same identifier. Here, we additionally leveraged WordNet to identify and replace sets of non-biomedical synonyms with their most common representatives. This combined approach aimed to reduce embedding noise from non-biomedical terms while preserving the integrity of biomedical concept representations. We applied this method to 1,055 biomedical concept sets representing molecular signatures or medical categories and assessed the mean pairwise distance of embeddings with and without non-biomedical synonym replacement. A smaller mean pairwise distance was interpreted as greater intra-cluster coherence and higher embedding quality. Embeddings were generated using the Word2Vec algorithm applied to a corpus of 10 million PubMed abstracts. Our results demonstrate that the addition of non-biomedical synonym replacement reduced the mean intra-cluster distance by an average of 8%, suggesting that this complementary approach enhances embedding quality. Future work will assess its applicability to other embedding techniques and downstream tasks. Python code implementing this method is provided under an open-source license.
Journal Article
Phenolyzer: phenotype-based prioritization of candidate genes for human diseases
2015
The Phenolyzer software provides prioritized candidate gene lists based on disease and phenotype terms that are entered by users as free text.
Prior biological knowledge and phenotype information may help to identify disease genes from human whole-genome and whole-exome sequencing studies. We developed Phenolyzer (
http://phenolyzer.usc.edu
), a tool that uses prior information to implicate genes involved in diseases. Phenolyzer exhibits superior performance over competing methods for prioritizing Mendelian and complex disease genes, based on disease or phenotype terms entered as free text.
Journal Article
E2F6 initiates stable epigenetic silencing of germline genes during embryonic development
2021
In mouse development, long-term silencing by CpG island DNA methylation is specifically targeted to germline genes; however, the molecular mechanisms of this specificity remain unclear. Here, we demonstrate that the transcription factor E2F6, a member of the polycomb repressive complex 1.6 (PRC1.6), is critical to target and initiate epigenetic silencing at germline genes in early embryogenesis. Genome-wide, E2F6 binds preferentially to CpG islands in embryonic cells. E2F6 cooperates with MGA to silence a subgroup of germline genes in mouse embryonic stem cells and in embryos, a function that critically depends on the E2F6 marked box domain. Inactivation of
E2f6
leads to a failure to deposit CpG island DNA methylation at these genes during implantation. Furthermore, E2F6 is required to initiate epigenetic silencing in early embryonic cells but becomes dispensable for the maintenance in differentiated cells. Our findings elucidate the mechanisms of epigenetic targeting of germline genes and provide a paradigm for how transient repression signals by DNA-binding factors in early embryonic cells are translated into long-term epigenetic silencing during mouse development.
DNA methylation targets CpG island promoters of germline genes to repress their expression in mouse somatic cells. Here the authors show that a transcription factor E2F6 is required to target CpG island DNA methylation and epigenetic silencing to germline genes during early mouse development.
Journal Article