Search Results Heading

MBRLSearchResults

mbrl.module.common.modules.added.book.to.shelf
Title added to your shelf!
View what I already have on My Shelf.
Oops! Something went wrong.
Oops! Something went wrong.
While trying to add the title to your shelf something went wrong :( Kindly try again later!
Are you sure you want to remove the book from the shelf?
Oops! Something went wrong.
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
    Done
    Filters
    Reset
  • Discipline
      Discipline
      Clear All
      Discipline
  • Is Peer Reviewed
      Is Peer Reviewed
      Clear All
      Is Peer Reviewed
  • Item Type
      Item Type
      Clear All
      Item Type
  • Subject
      Subject
      Clear All
      Subject
  • Year
      Year
      Clear All
      From:
      -
      To:
  • More Filters
      More Filters
      Clear All
      More Filters
      Source
    • Language
802 result(s) for "sequence coevolution"
Sort by:
Nuclear Egress Complexes of HCMV and Other Herpesviruses: Solving the Puzzle of Sequence Coevolution, Conserved Structures and Subfamily-Spanning Binding Properties
Herpesviruses uniquely express two essential nuclear egress-regulating proteins forming a heterodimeric nuclear egress complex (core NEC). These core NECs serve as hexameric lattice-structured platforms for capsid docking and recruit viral and cellular NEC-associated factors that jointly exert nuclear lamina as well as membrane-rearranging functions (multicomponent NEC). The regulation of nuclear egress has been profoundly analyzed for murine and human cytomegaloviruses (CMVs) on a mechanistic basis, followed by the description of core NEC crystal structures, first for HCMV, then HSV-1, PRV and EBV. Interestingly, the highly conserved structural domains of these proteins stand in contrast to a very limited sequence conservation of the key amino acids within core NEC-binding interfaces. Even more surprising, although a high functional consistency was found when regarding the basic role of NECs in nuclear egress, a clear specification was identified regarding the limited, subfamily-spanning binding properties of core NEC pairs and NEC multicomponent proteins. This review summarizes the evolving picture of the relationship between sequence coevolution, structural conservation and properties of NEC interaction, comparing HCMV to α-, β- and γ-herpesviruses. Since NECs represent substantially important elements of herpesviral replication that are considered as drug-accessible targets, their putative translational use for antiviral strategies is discussed.
cpxDeepMSA: A Deep Cascade Algorithm for Constructing Multiple Sequence Alignments of Protein–Protein Interactions
Protein–protein interactions (PPIs) are fundamental to many biological processes. The coevolution-based prediction of interacting residues has made great strides in protein complexes that are known to interact. A multiple sequence alignment (MSA) is the basis of coevolution analysis. MSAs have recently made significant progress in the protein monomer sequence analysis. However, no standard or efficient pipelines are available for the sensitive protein complex MSA (cpxMSA) collection. How to generate cpxMSA is one of the most challenging problems of sequence coevolution analysis. Although several methods have been developed to address this problem, no standalone program exists. Furthermore, the number of built-in properties is limited; hence, it is often difficult for users to analyze sequence coevolution according to their desired cpxMSA. In this article, we developed a novel cpxMSA approach (cpxDeepMSA. We used different protein monomer databases and incorporated the three strategies (genomic distance, phylogeny information, and STRING interaction network) used to join the monomer MSA results of protein complexes, which can prevent using a single method fail to the joint two-monomer MSA causing the cpxMSA construction failure. We anticipate that the cpxDeepMSA algorithm will become a useful high-throughput tool in protein complex structure predictions, inter-protein residue-residue contacts, and the biological sequence coevolution analysis.
Antiparallel protocadherin homodimers use distinct affinity- and specificity-mediating regions in cadherin repeats 1-4
Protocadherins (Pcdhs) are cell adhesion and signaling proteins used by neurons to develop and maintain neuronal networks, relying on trans homophilic interactions between their extracellular cadherin (EC) repeat domains. We present the structure of the antiparallel EC1-4 homodimer of human PcdhγB3, a member of the γ subfamily of clustered Pcdhs. Structure and sequence comparisons of α, β, and γ clustered Pcdh isoforms illustrate that subfamilies encode specificity in distinct ways through diversification of loop region structure and composition in EC2 and EC3, which contains isoform-specific conservation of primarily polar residues. In contrast, the EC1/EC4 interface comprises hydrophobic interactions that provide non-selective dimerization affinity. Using sequence coevolution analysis, we found evidence for a similar antiparallel EC1-4 interaction in non-clustered Pcdh families. We thus deduce that the EC1-4 antiparallel homodimer is a general interaction strategy that evolved before the divergence of these distinct protocadherin families. As the brain develops, nerve cells or neurons connect with one another to form complex networks. These connections form between branch-like structures, called dendrites, that project from the cell body of each neuron. To prevent unneeded connections from forming, dendrites that belong to the same neuron need a way to recognize and avoid one another. A family of proteins called protocadherins supports this process of self-avoidance. Protocadherins have three main parts or domains: an extracellular domain that faces outwards away from the cell, a transmembrane domain that sits within the cell’s surface membrane and an intracellular domain that faces into the cell’s interior. There are two major groups of protocadherins – clustered and non-clustered – and the former are responsible for the self-avoidance behavior between dendrites. Clustered protocadherins in turn comprise three subfamilies, each of which consists of multiple variants with slightly different structures (known as isoforms). The particular set of protocadherin isoforms that a neuron displays on its surface distinguishes that neuron from all others, a little like a barcode. When two dendrites meet, the protocadherins in their membranes come into contact with one another. If both dendrites come from the same neuron and therefore possess identical sets of protocadherins, then all protocadherins can form two-subunit complexes containing one copy of the same isoform from each dendrite. These complexes are called homodimers and their formation acts as a signal that informs the cell that it has encountered one of its own dendrites and should therefore not establish a connection. By using X-rays to determine the structure of a crystallized protocadherin fragment down to the level of its individual atoms, Nicoludis et al. now reveal exactly how clustered protocadherins form homodimers. The results show that each protocadherin subfamily uses a slightly different type of interaction due to differences in the structure of their extracellular domains. The next challenge is to identify the signaling cascade that is triggered by the formation of clustered protocadherin homodimers, and to work out how activation of this cascade prevents a permanent connection from forming. In addition, the results of Nicoludis et al. predict that some non-clustered protocadherins form dimers with a similar architecture to that of clustered protocadherins. This possibility should also be tested experimentally.
Gram-negative outer-membrane proteins with multiple β-barrel domains
Outer-membrane beta barrels (OMBBs) are found in the outer membrane of gram-negative bacteria and eukaryotic organelles. OMBBs fold as antiparallel β-sheets that close onto themselves, forming pores that traverse the membrane. Currently known structures include only one barrel, of 8 to 36 strands, per chain. The lack of multi-OMBB chains is surprising, as most OMBBs form oligomers, and some function only in this state. Using a combination of sensitive sequence comparison methods and coevolutionary analysis tools, we identify many proteins combining multiple beta barrels within a single chain; combinations that include eight-stranded barrels prevail. These multibarrels seem to be the result of independent, lineage-specific fusion and amplification events. The absence of multibarrels that are universally conserved in bacteria with an outer membrane, coupled with their frequent de novo genesis, suggests that their functions are not essential but rather beneficial in specific environments. Adjacent barrels of complementary function within the same chain may allow for functions beyond those of the individual barrels.
Assessing the utility of coevolution-based residue–residue contact predictions in a sequence- and structure-rich era
Recently developed methods have shown considerable promise in predicting residue–residue contacts in protein 3D structures using evolutionary covariance information. However, these methods require large numbers of evolutionarily related sequences to robustly assess the extent of residue covariation, and the larger the protein family, the more likely that contact information is unnecessary because a reasonable model can be built based on the structure of a homolog. Here we describe a method that integrates sequence coevolution and structural context information using a pseudolikelihood approach, allowing more accurate contact predictions from fewer homologous sequences. We rigorously assess the utility of predicted contacts for protein structure prediction using large and representative sequence and structure databases from recent structure prediction experiments. We find that contact predictions are likely to be accurate when the number of aligned sequences (with sequence redundancy reduced to 90%) is greater than five times the length of the protein, and that accurate predictions are likely to be useful for structure modeling if the aligned sequences are more similar to the protein of interest than to the closest homolog of known structure. These conditions are currently met by 422 of the protein families collected in the Pfam database.
Deep learning suggests that gene expression is encoded in all parts of a co-evolving interacting gene regulatory structure
Understanding the genetic regulatory code governing gene expression is an important challenge in molecular biology. However, how individual coding and non-coding regions of the gene regulatory structure interact and contribute to mRNA expression levels remains unclear. Here we apply deep learning on over 20,000 mRNA datasets to examine the genetic regulatory code controlling mRNA abundance in 7 model organisms ranging from bacteria to Human. In all organisms, we can predict mRNA abundance directly from DNA sequence, with up to 82% of the variation of transcript levels encoded in the gene regulatory structure. By searching for DNA regulatory motifs across the gene regulatory structure, we discover that motif interactions could explain the whole dynamic range of mRNA levels. Co-evolution across coding and non-coding regions suggests that it is not single motifs or regions, but the entire gene regulatory structure and specific combination of regulatory elements that define gene expression levels. Regulatory and coding regions of genes are shaped by evolution to control expression levels. Here, the authors use deep learning to identify rules controlling gene expression levels and suggest that all parts of the gene regulatory structure interact in this.
ProteinNet: a standardized data set for machine learning of protein structure
Background Rapid progress in deep learning has spurred its application to bioinformatics problems including protein structure prediction and design. In classic machine learning problems like computer vision, progress has been driven by standardized data sets that facilitate fair assessment of new methods and lower the barrier to entry for non-domain experts. While data sets of protein sequence and structure exist, they lack certain components critical for machine learning, including high-quality multiple sequence alignments and insulated training/validation splits that account for deep but only weakly detectable homology across protein space. Results We created the ProteinNet series of data sets to provide a standardized mechanism for training and assessing data-driven models of protein sequence-structure relationships. ProteinNet integrates sequence, structure, and evolutionary information in programmatically accessible file formats tailored for machine learning frameworks. Multiple sequence alignments of all structurally characterized proteins were created using substantial high-performance computing resources. Standardized data splits were also generated to emulate the difficulty of past CASP (Critical Assessment of protein Structure Prediction) experiments by resetting protein sequence and structure space to the historical states that preceded six prior CASPs. Utilizing sensitive evolution-based distance metrics to segregate distantly related proteins, we have additionally created validation sets distinct from the official CASP sets that faithfully mimic their difficulty. Conclusion ProteinNet represents a comprehensive and accessible resource for training and assessing machine-learned models of protein structure.
Conformational buffering underlies functional selection in intrinsically disordered protein regions
Many disordered proteins conserve essential functions in the face of extensive sequence variation, making it challenging to identify the mechanisms responsible for functional selection. Here we identify the molecular mechanism of functional selection for the disordered adenovirus early gene 1A (E1A) protein. E1A competes with host factors to bind the retinoblastoma (Rb) protein, subverting cell cycle regulation. We show that two binding motifs tethered by a hypervariable disordered linker drive picomolar affinity Rb binding and host factor displacement. Compensatory changes in amino acid sequence composition and sequence length lead to conservation of optimal tethering across a large family of E1A linkers. We refer to this compensatory mechanism as conformational buffering. We also detect coevolution of the motifs and linker, which can preserve or eliminate the tethering mechanism. Conformational buffering and motif–linker coevolution explain robust functional encoding within hypervariable disordered linkers and could underlie functional selection of many disordered protein regions. Foutel et. al. identify conformational buffering as a mechanism for functional selection in intrinsically disordered protein regions that allows robust encoding of a tethering function by a hypervariable disordered linker through compensatory changes in sequence length and composition.
CopulaNet: Learning residue co-evolution directly from multiple sequence alignment for protein structure prediction
Residue co-evolution has become the primary principle for estimating inter-residue distances of a protein, which are crucially important for predicting protein structure. Most existing approaches adopt an indirect strategy, i.e., inferring residue co-evolution based on some hand-crafted features, say, a covariance matrix, calculated from multiple sequence alignment (MSA) of target protein. This indirect strategy, however, cannot fully exploit the information carried by MSA. Here, we report an end-to-end deep neural network, CopulaNet, to estimate residue co-evolution directly from MSA. The key elements of CopulaNet include: (i) an encoder to model context-specific mutation for each residue; (ii) an aggregator to model residue co-evolution, and thereafter estimate inter-residue distances. Using CASP13 (the 13th Critical Assessment of Protein Structure Prediction) target proteins as representatives, we demonstrate that CopulaNet can predict protein structure with improved accuracy and efficiency. This study represents a step toward improved end-to-end prediction of inter-residue distances and protein tertiary structures. Protein structure prediction is a challenge. A new deep learning framework, CopulaNet, is a major step forward toward end-to-end prediction of inter-residue distances and protein tertiary structures with improved accuracy and efficiency.
Molecular evolution of peptidergic signaling systems in bilaterians
Peptide hormones and their receptors are widespread in metazoans, but the knowledge we have of their evolutionary relationships remains unclear. Recently, accumulating genome sequences from many different species have offered the opportunity to reassess the relationships between protostomian and deuterostomian peptidergic systems (PSs). Here we used sequences of all human rhodopsin and secretin-type G protein-coupled receptors as bait to retrieve potential homologs in the genomes of 15 bilaterian species, including nonchordate deuterostomian and lophotrochozoan species. Our phylogenetic analysis of these receptors revealed 29 well-supported subtrees containing mixed sets of protostomian and deuterostomian sequences. This indicated that many vertebrate and arthropod PSs that were previously thought to be phyla specific are in fact of bilaterian origin. By screening sequence databases for potential peptides, we then reconstructed entire bilaterian peptide families and showed that protostomian and deuterostomian peptides that are ligands of orthologous receptors displayed some similarity at the level of their primary sequence, suggesting an ancient coevolution between peptide and receptor genes. In addition to shedding light on the function of human G protein-coupled receptor PSs, this work presents orthology markers to study ancestral neuron types that were probably present in the last common bilaterian ancestor.