Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
247,912
result(s) for
"protein models"
Sort by:
Introduction to protein structure prediction
by
Rangwala, Huzefa
,
Karypis, George
in
Bioinformatics & Computational Biology
,
Computer simulation
,
Life Sciences
2010
A look at the methods and algorithms used to predict protein structure A thorough knowledge of the function and structure of proteins is critical for the advancement of biology and the life sciences as well as the development of better drugs, higher-yield crops, and even synthetic bio-fuels. To that end, this reference sheds light on the methods used for protein structure prediction and reveals the key applications of modeled structures. This indispensable book covers the applications of modeled protein structures and unravels the relationship between pure sequence information and three-dimensional structure, which continues to be one of the greatest challenges in molecular biology. With this resource, readers will find an all-encompassing examination of the problems, methods, tools, servers, databases, and applications of protein structure prediction and they will acquire unique insight into the future applications of the modeled protein structures. The book begins with a thorough introduction to the protein structure prediction problem and is divided into four themes: a background on structure prediction, the prediction of structural elements, tertiary structure prediction, and functional insights. Within those four sections, the following topics are covered: Databases and resources that are commonly used for protein structure prediction The structure prediction flagship assessment (CASP) and the protein structure initiative (PSI) Definitions of recurring substructures and the computational approaches used for solving sequence problems Difficulties with contact map prediction and how sophisticated machine learning methods can solve those problems Structure prediction methods that rely on homology modeling, threading, and fragment assembly Hybrid methods that achieve high-resolution protein structures Parts of the protein structure that may be conserved and used to interact with other biomolecules How the loop prediction problem can be used for refinement of the modeled structures The computational model that detects the differences between protein structure and its modeled mutant Whether working in the field of bioinformatics or molecular biology research or taking courses in protein modeling, readers will find the content in this book invaluable.
DisoFLAG: accurate prediction of protein intrinsic disorder and its functions using graph-based interaction protein language model
2024
Intrinsically disordered proteins and regions (IDPs/IDRs) are functionally important proteins and regions that exist as highly dynamic conformations under natural physiological conditions. IDPs/IDRs exhibit a broad range of molecular functions, and their functions involve binding interactions with partners and remaining native structural flexibility. The rapid increase in the number of proteins in sequence databases and the diversity of disordered functions challenge existing computational methods for predicting protein intrinsic disorder and disordered functions. A disordered region interacts with different partners to perform multiple functions, and these disordered functions exhibit different dependencies and correlations. In this study, we introduce DisoFLAG, a computational method that leverages a graph-based interaction protein language model (GiPLM) for jointly predicting disorder and its multiple potential functions. GiPLM integrates protein semantic information based on pre-trained protein language models into graph-based interaction units to enhance the correlation of the semantic representation of multiple disordered functions. The DisoFLAG predictor takes amino acid sequences as the only inputs and provides predictions of intrinsic disorder and six disordered functions for proteins, including protein-binding, DNA-binding, RNA-binding, ion-binding, lipid-binding, and flexible linker. We evaluated the predictive performance of DisoFLAG following the Critical Assessment of protein Intrinsic Disorder (CAID) experiments, and the results demonstrated that DisoFLAG offers accurate and comprehensive predictions of disordered functions, extending the current coverage of computationally predicted disordered function categories. The standalone package and web server of DisoFLAG have been established to provide accurate prediction tools for intrinsic disorders and their associated functions.
Journal Article
Designing and evaluating the MULTICOM protein local and global model quality prediction methods in the CASP10 experiment
by
Wang, Zheng
,
Cao, Renzhi
,
Cheng, Jianlin
in
Biochemistry
,
Biomedical and Life Sciences
,
Computational analysis
2014
Background
Protein model quality assessment is an essential component of generating and using protein structural models. During the Tenth Critical Assessment of Techniques for Protein Structure Prediction (CASP10), we developed and tested four automated methods (MULTICOM-REFINE, MULTICOM-CLUSTER, MULTICOM-NOVEL, and MULTICOM-CONSTRUCT) that predicted both local and global quality of protein structural models.
Results
MULTICOM-REFINE was a clustering approach that used the average pairwise structural similarity between models to measure the global quality and the average Euclidean distance between a model and several top ranked models to measure the local quality. MULTICOM-CLUSTER and MULTICOM-NOVEL were two new support vector machine-based methods of predicting both the local and global quality of a single protein model. MULTICOM-CONSTRUCT was a new weighted pairwise model comparison (clustering) method that used the weighted average similarity between models in a pool to measure the global model quality. Our experiments showed that the pairwise model assessment methods worked better when a large portion of models in the pool were of good quality, whereas single-model quality assessment methods performed better on some hard targets when only a small portion of models in the pool were of reasonable quality.
Conclusions
Since digging out a few good models from a large pool of low-quality models is a major challenge in protein structure prediction, single model quality assessment methods appear to be poised to make important contributions to protein structure modeling. The other interesting finding was that single-model quality assessment scores could be used to weight the models by the consensus pairwise model comparison method to improve its accuracy.
Journal Article
Structure, function and regulation of the hsp90 machinery
by
Buchner, Johannes
,
Li, Jing
in
85747 Garching Germany Login to access the Email id Crossref citations 19 PMC citations 11 DOI: 10.4103/2319-4170.113230 PMID: 23806880 Get Permissions Abstract Heat shock protein 90 (Hsp90) is an ATP-dependent molecular chaperone which is essential in eukaryotes. It is required for the activation and stabilization of a wide variety of client proteins and many of them are involved in important cellular pathways. Since Hsp90 affects numerous physiological processes such as signal transduction
,
a middle domain (M-domain)
,
a new model of the chaperone cycle emerges [Figure 3]A
2013
Heat shock protein 90 (Hsp90) is an ATP-dependent molecular chaperone which is essential in eukaryotes. It is required for the activation and stabilization of a wide variety of client proteins and many of them are involved in important cellular pathways. Since Hsp90 affects numerous physiological processes such as signal transduction, intracellular transport, and protein degradation, it became an interesting target for cancer therapy. Structurally, Hsp90 is a flexible dimeric protein composed of three different domains which adopt structurally distinct conformations. ATP binding triggers directionality in these conformational changes and leads to a more compact state. To achieve its function, Hsp90 works together with a large group of cofactors, termed co-chaperones. Co-chaperones form defined binary or ternary complexes with Hsp90, which facilitate the maturation of client proteins. In addition, posttranslational modifications of Hsp90, such as phosphorylation and acetylation, provide another level of regulation. They influence the conformational cycle, co-chaperone interaction, and inter-domain communications. In this review, we discuss the recent progress made in understanding the Hsp90 machinery.
Journal Article
SAMMSON fosters cancer cell fitness by concertedly enhancing mitochondrial and cytosolic translation
by
Vendramin, Roberto
,
Lafontaine, Denis L J
,
Saraf, Kritika
in
Cancer
,
Complex formation
,
Cytosol
2018
Synchronization of mitochondrial and cytoplasmic translation rates is critical for the maintenance of cellular fitness, with cancer cells being especially vulnerable to translational uncoupling. Although alterations of cytosolic protein synthesis are common in human cancer, compensating mechanisms in mitochondrial translation remain elusive. Here we show that the malignant long non-coding RNA (lncRNA) SAMMSON promotes a balanced increase in ribosomal RNA (rRNA) maturation and protein synthesis in the cytosol and mitochondria by modulating the localization of CARF, an RNA-binding protein that sequesters the exo-ribonuclease XRN2 in the nucleoplasm, which under normal circumstances limits nucleolar rRNA maturation. SAMMSON interferes with XRN2 binding to CARF in the nucleus by favoring the formation of an aberrant cytoplasmic RNA–protein complex containing CARF and p32, a mitochondrial protein required for the processing of the mitochondrial rRNAs. These data highlight how a single oncogenic lncRNA can simultaneously modulate RNA–protein complex formation in two distinct cellular compartments to promote cell growth.
Journal Article
DeepQA: improving the estimation of single protein model quality with deep belief networks
by
Bhattacharya, Debswapna
,
Cao, Renzhi
,
Cheng, Jianlin
in
Algorithms
,
Bioinformatics
,
Biomedical and Life Sciences
2016
Background
Protein quality assessment (QA) useful for ranking and selecting protein models has long been viewed as one of the major challenges for protein tertiary structure prediction. Especially, estimating the quality of a single protein model, which is important for selecting a few good models out of a large model pool consisting of mostly low-quality models, is still a largely unsolved problem.
Results
We introduce a novel single-model quality assessment method DeepQA based on deep belief network that utilizes a number of selected features describing the quality of a model from different perspectives, such as energy, physio-chemical characteristics, and structural information. The deep belief network is trained on several large datasets consisting of models from the Critical Assessment of Protein Structure Prediction (CASP) experiments, several publicly available datasets, and models generated by our in-house
ab initio
method. Our experiments demonstrate that deep belief network has better performance compared to Support Vector Machines and Neural Networks on the protein model quality assessment problem, and our method DeepQA achieves the state-of-the-art performance on CASP11 dataset. It also outperformed two well-established methods in selecting good outlier models from a large set of models of mostly low quality generated by
ab initio
modeling methods.
Conclusion
DeepQA is a useful deep learning tool for protein single model quality assessment and protein structure prediction. The source code, executable, document and training/test datasets of DeepQA for Linux is freely available to non-commercial users at
http://cactus.rnet.missouri.edu/DeepQA/
.
Journal Article
TMbed: transmembrane proteins predicted through language model embeddings
2022
Background
Despite the immense importance of transmembrane proteins (TMP) for molecular biology and medicine, experimental 3D structures for TMPs remain about 4–5 times underrepresented compared to non-TMPs. Today’s top methods such as AlphaFold2 accurately predict 3D structures for many TMPs, but annotating transmembrane regions remains a limiting step for proteome-wide predictions.
Results
Here, we present TMbed, a novel method inputting embeddings from protein Language Models (pLMs, here ProtT5), to predict for each residue one of four classes: transmembrane helix (TMH), transmembrane strand (TMB), signal peptide, or other. TMbed completes predictions for entire proteomes within hours on a single consumer-grade desktop machine at performance levels similar or better than methods, which are using evolutionary information from multiple sequence alignments (MSAs) of protein families. On the per-protein level, TMbed correctly identified 94 ± 8% of the beta barrel TMPs (53 of 57) and 98 ± 1% of the alpha helical TMPs (557 of 571) in a non-redundant data set, at false positive rates well below 1% (erred on 30 of 5654 non-membrane proteins). On the per-segment level, TMbed correctly placed, on average, 9 of 10 transmembrane segments within five residues of the experimental observation. Our method can handle sequences of up to 4200 residues on standard graphics cards used in desktop PCs (e.g., NVIDIA GeForce RTX 3060).
Conclusions
Based on embeddings from pLMs and two novel filters (Gaussian and Viterbi), TMbed predicts alpha helical and beta barrel TMPs at least as accurately as any other method but at lower false positive rates. Given the few false positives and its outstanding speed, TMbed might be ideal to sieve through millions of 3D structures soon to be predicted, e.g., by AlphaFold2.
Journal Article
Improved model quality assessment using ProQ2
2012
Background
Employing methods to assess the quality of modeled protein structures is now standard practice in bioinformatics. In a broad sense, the techniques can be divided into methods relying on consensus prediction on the one hand, and
single-model
methods on the other. Consensus methods frequently perform very well when there is a clear consensus, but this is not always the case. In particular, they frequently fail in selecting the best possible model in the hard cases (lacking consensus) or in the easy cases where models are very similar. In contrast, single-model methods do not suffer from these drawbacks and could potentially be applied on any protein of interest to assess quality or as a scoring function for sampling-based refinement.
Results
Here, we present a new single-model method, ProQ2, based on ideas from its predecessor, ProQ. ProQ2 is a model quality assessment algorithm that uses support vector machines to predict local as well as global quality of protein models. Improved performance is obtained by combining previously used features with updated structural and predicted features. The most important contribution can be attributed to the use of profile weighting of the residue specific features and the use features averaged over the whole model even though the prediction is still local.
Conclusions
ProQ2 is significantly better than its predecessors at detecting high quality models, improving the sum of Z-scores for the selected first-ranked models by 20% and 32% compared to the second-best single-model method in CASP8 and CASP9, respectively. The absolute quality assessment of the models at both local and global level is also improved. The Pearson’s correlation between the correct and local predicted score is improved from 0.59 to 0.70 on CASP8 and from 0.62 to 0.68 on CASP9; for global score to the correct GDT_TS from 0.75 to 0.80 and from 0.77 to 0.80 again compared to the second-best single methods in CASP8 and CASP9, respectively. ProQ2 is available at
http://proq2.wallnerlab.org
.
Journal Article
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences
by
Goyal, Siddharth
,
Ma, Jerry
,
Guo, Demi
in
Amino acid sequence
,
Amino acids
,
Artificial intelligence
2021
In the field of artificial intelligence, a combination of scale in data and model capacity enabled by unsupervised learning has led to major advances in representation learning and statistical generation. In the life sciences, the anticipated growth of sequencing promises unprecedented data on natural sequence diversity. Protein language modeling at the scale of evolution is a logical step toward predictive and generative artificial intelligence for biology. To this end, we use unsupervised learning to train a deep contextual language model on 86 billion amino acids across 250 million protein sequences spanning evolutionary diversity. The resulting model contains information about biological properties in its representations. The representations are learned from sequence data alone. The learned representation space has a multiscale organization reflecting structure from the level of biochemical properties of amino acids to remote homology of proteins. Information about secondary and tertiary structure is encoded in the representations and can be identified by linear projections. Representation learning produces features that generalize across a range of applications, enabling state-of-the-art supervised prediction of mutational effect and secondary structure and improving state-of-the-art features for long-range contact prediction.
Journal Article
Is Protein Folding a Thermodynamically Unfavorable, Active, Energy-Dependent Process?
2022
The prevailing current view of protein folding is the thermodynamic hypothesis, under which the native folded conformation of a protein corresponds to the global minimum of Gibbs free energy G. We question this concept and show that the empirical evidence behind the thermodynamic hypothesis of folding is far from strong. Furthermore, physical theory-based approaches to the prediction of protein folds and their folding pathways so far have invariably failed except for some very small proteins, despite decades of intensive theory development and the enormous increase of computer power. The recent spectacular successes in protein structure prediction owe to evolutionary modeling of amino acid sequence substitutions enhanced by deep learning methods, but even these breakthroughs provide no information on the protein folding mechanisms and pathways. We discuss an alternative view of protein folding, under which the native state of most proteins does not occupy the global free energy minimum, but rather, a local minimum on a fluctuating free energy landscape. We further argue that ΔG of folding is likely to be positive for the majority of proteins, which therefore fold into their native conformations only through interactions with the energy-dependent molecular machinery of living cells, in particular, the translation system and chaperones. Accordingly, protein folding should be modeled as it occurs in vivo, that is, as a non-equilibrium, active, energy-dependent process.
Journal Article