Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
19
result(s) for
"Findeiß, Sven"
Sort by:
Proteinortho: Detection of (Co-)orthologs in large-scale analysis
by
Prohaska, Sonja J
,
Steiner, Lydia
,
Lechner, Marcus
in
Algorithms
,
Applications software
,
Base Sequence
2011
Background
Orthology analysis is an important part of data analysis in many areas of bioinformatics such as comparative genomics and molecular phylogenetics. The ever-increasing flood of sequence data, and hence the rapidly increasing number of genomes that can be compared simultaneously, calls for efficient software tools as brute-force approaches with quadratic memory requirements become infeasible in practise. The rapid pace at which new data become available, furthermore, makes it desirable to compute genome-wide orthology relations for a given dataset rather than relying on relations listed in databases.
Results
The program
Proteinortho
described here is a stand-alone tool that is geared towards large datasets and makes use of distributed computing techniques when run on multi-core hardware. It implements an extended version of the reciprocal best alignment heuristic. We apply
Proteinortho
to compute orthologous proteins in the complete set of all 717 eubacterial genomes available at NCBI at the beginning of 2009. We identified thirty proteins present in 99% of all bacterial proteomes.
Conclusions
Proteinortho
significantly reduces the required amount of memory for orthology analysis compared to existing tools, allowing such computations to be performed on off-the-shelf hardware.
Journal Article
The primary transcriptome of the major human pathogen Helicobacter pylori
by
Vogel, Jörg
,
Reignier, Jérémy
,
Sharma, Cynthia M.
in
5' Untranslated Regions - genetics
,
631/208/212/2019
,
631/208/514/1949
2010
Genome sequencing of
Helicobacter pylori
has revealed the potential proteins and genetic diversity of this prevalent human pathogen, yet little is known about its transcriptional organization and noncoding RNA output. Massively parallel cDNA sequencing (RNA-seq) has been revolutionizing global transcriptomic analysis. Here, using a novel differential approach (dRNA-seq) selective for the 5′ end of primary transcripts, we present a genome-wide map of
H. pylori
transcriptional start sites and operons. We discovered hundreds of transcriptional start sites within operons, and opposite to annotated genes, indicating that complexity of gene expression from the small
H. pylori
genome is increased by uncoupling of polycistrons and by genome-wide antisense transcription. We also discovered an unexpected number of ∼60 small RNAs including the ε-subdivision counterpart of the regulatory 6S RNA and associated RNA products, and potential regulators of
cis
- and
trans
-encoded target messenger RNAs. Our approach establishes a paradigm for mapping and annotating the primary transcriptomes of many living species.
The
Helicobacter pylori
transcriptome
The bacterium
Helicobacter pylori
infects about half the human population, thriving in the acid conditions of the stomach. Most carriers are asymptomatic, but some suffer inflammation, ulcers and gastric cancer. Now, using a novel approach that selects for the 5' ends of primary transcripts, the 'primary transcriptome' — predominantly the unprocessed messenger RNAs and small non-coding RNAs — has been determined for
H. pylori
in a variety of growth conditions. With the genome sequence and protein interactome previously published, this work provides the third global reference data set for the widely used
Helicobacter
strain 26695.
The transcriptome of
Helicobacter pylori
, an important human pathogen involved in gastric ulcers and cancer, is presented. The approach establishes a model for mapping and annotating the primary transcriptomes of many living species.
Journal Article
Design of Artificial Riboswitches as Biosensors
by
Mörl, Mario
,
Stadler, Peter
,
Findeiß, Sven
in
aptamer
,
Aptamers, Nucleotide
,
Biosensing Techniques
2017
RNA aptamers readily recognize small organic molecules, polypeptides, as well as other nucleic acids in a highly specific manner. Many such aptamers have evolved as parts of regulatory systems in nature. Experimental selection techniques such as SELEX have been very successful in finding artificial aptamers for a wide variety of natural and synthetic ligands. Changes in structure and/or stability of aptamers upon ligand binding can propagate through larger RNA constructs and cause specific structural changes at distal positions. In turn, these may affect transcription, translation, splicing, or binding events. The RNA secondary structure model realistically describes both thermodynamic and kinetic aspects of RNA structure formation and refolding at a single, consistent level of modelling. Thus, this framework allows studying the function of natural riboswitches in silico. Moreover, it enables rationally designing artificial switches, combining essentially arbitrary sensors with a broad choice of read-out systems. Eventually, this approach sets the stage for constructing versatile biosensors.
Journal Article
Computational Analysis of Telomerase RNA Evolution in Caenorhabditis Species
by
Reinhardt, Franziska
,
Stadler, Peter F.
,
Klapproth, Christopher
in
Analysis
,
Annotations
,
Arthropods
2026
Background/Objectives: The telomerase RNA (TR) is an indispensable part of the telomerase protein complex responsible for telomere elongation in most eukaryotic species. Although the telomere terminal repeat sequence (TTAGGC)n in Caenorhabditis elegans has been known for years, a telomerase RNA gene was not identified in the entire phylum of Nematoda until recently. Methods: In this exploratory study, we employ a combination of different approaches to identify likely telomerase RNA candidates among putative non-coding transcripts. Results: A detailed analysis of our prime candidate shows compelling evidence that it encodes the missing RNA element of the telomerase complex, which is notably located in an intron of the coding gene nmy-2. Using nmy-2 homologs in other nematodes as anchors, we annotate the conserved TR gene in 21 Caenorhabditis species. We furthermore show that the intronic localization of the TR gene is conserved in two distinct branching groups of the Caenorhabditis phylogeny and demonstrate that this property likely emerged from a single point of origin. Conclusions: While the intronic TR represents a very interesting evolutionary adaption that seems to have been successful in the Elegans and Japonica groups, the question regarding the macroscopic nematode TR evolution remains.
Journal Article
Common Features in lncRNA Annotation and Classification: A Survey
by
Stadler, Peter F.
,
Fallmann, Jörg
,
Klapproth, Christopher
in
Amino acids
,
Annotations
,
Chromatin
2021
Long non-coding RNAs (lncRNAs) are widely recognized as important regulators of gene expression. Their molecular functions range from miRNA sponging to chromatin-associated mechanisms, leading to effects in disease progression and establishing them as diagnostic and therapeutic targets. Still, only a few representatives of this diverse class of RNAs are well studied, while the vast majority is poorly described beyond the existence of their transcripts. In this review we survey common in silico approaches for lncRNA annotation. We focus on the well-established sets of features used for classification and discuss their specific advantages and weaknesses. While the available tools perform very well for the task of distinguishing coding sequence from other RNAs, we find that current methods are not well suited to distinguish lncRNAs or parts thereof from other non-protein-coding input sequences. We conclude that the distinction of lncRNAs from intronic sequences and untranslated regions of coding mRNAs remains a pressing research gap.
Journal Article
TSSAR: TSS annotation regime for dRNA-seq data
2014
Background
Differential RNA sequencing (dRNA-seq) is a high-throughput screening technique designed to examine the architecture of bacterial operons in general and the precise position of transcription start sites (TSS) in particular. Hitherto, dRNA-seq data were analyzed by visualizing the sequencing reads mapped to the reference genome and manually annotating reliable positions. This is very labor intensive and, due to the subjectivity, biased.
Results
Here, we present
TSSAR
, a tool for automated
de novo
TSS annotation from dRNA-seq data that respects the statistics of dRNA-seq libraries.
TSSAR
uses the premise that the number of sequencing reads starting at a certain genomic position within a transcriptional active region follows a Poisson distribution with a parameter that depends on the local strength of expression. The differences of two dRNA-seq library counts thus follow a Skellam distribution. This provides a statistical basis to identify significantly enriched primary transcripts.
We assessed the performance by analyzing a publicly available dRNA-seq data set using
TSSAR
and two simple approaches that utilize user-defined score cutoffs. We evaluated the power of reproducing the manual TSS annotation. Furthermore, the same data set was used to reproduce 74 experimentally validated TSS in
H. pylori
from reliable techniques such as RACE or primer extension. Both analyses showed that
TSSAR
outperforms the static cutoff-dependent approaches.
Conclusions
Having an automated and efficient tool for analyzing dRNA-seq data facilitates the use of the dRNA-seq technique and promotes its application to more sophisticated analysis. For instance, monitoring the plasticity and dynamics of the transcriptomal architecture triggered by different stimuli and growth conditions becomes possible.
The main asset of a novel tool for dRNA-seq analysis that reaches out to a broad user community is usability. As such, we provide
TSSAR
both as intuitive RESTful Web service (
http://rna.tbi.univie.ac.at/TSSAR
) together with a set of post-processing and analysis tools, as well as a stand-alone version for use in high-throughput dRNA-seq data analysis pipelines.
Journal Article
small RNA Aar in Acinetobacter baylyi: a putative regulator of amino acid metabolism
by
Gerischer, Ulrike
,
Schilling, Dominik
,
Richter, Andreas S
in
A. baylyi
,
Acinetobacter
,
Acinetobacter - genetics
2010
Small non-coding RNAs (sRNAs) are key players in prokaryotic metabolic circuits, allowing the cell to adapt to changing environmental conditions. Regulatory interference by sRNAs in cellular metabolism is often facilitated by the Sm-like protein Hfq. A search for novel sRNAs in A. baylyi intergenic regions was performed by a biocomputational screening. One candidate, Aar, encoded between trpS and sucD showed Hfq dependency in Northern blot analysis. Aar was expressed strongly during stationary growth phase in minimal medium; in contrast, in complex medium, strongest expression was in the exponential growth phase. Whereas over-expression of Aar in trans did not affect bacterial growth, seven mRNA targets predicted by two in silico approaches were upregulated in stationary growth phase. All seven mRNAs are involved in A. baylyi amino acid metabolism. A putative binding site for Lrp, the global regulator of branched-chain amino acids in E. coli, was observed within the aar gene. Both facts imply an Aar participation in amino acid metabolism.
Journal Article
Tailored machine learning models for functional RNA detection in genome-wide screens
by
Fallmann, Jörg
,
Klapproth, Christopher
,
Kühnl, Felix
in
Annotations
,
Artificial intelligence
,
Bioinformatics
2023
The in silico prediction of non-coding and protein-coding genetic loci has received considerable attention in comparative genomics aiming in particular at the identification of properties of nucleotide sequences that are informative of their biological role in the cell. We present here a software framework for the alignment-based training, evaluation and application of machine learning models with user-defined parameters. Instead of focusing on the one-size-fits-all approach of pervasive in silico annotation pipelines, we offer a framework for the structured generation and evaluation of models based on arbitrary features and input data, focusing on stable and explainable results. Furthermore, we showcase the usage of our software package in a full-genome screen of Drosophila melanogaster and evaluate our results against the well-known but much less flexible program RNAz.
Journal Article
Optimization of parameters for coverage of low molecular weight proteins
by
Kalkhof, Stefan
,
Kohajda, Tibor
,
Kellis, Manolis
in
Analytical Chemistry
,
Annotations
,
Biochemistry
2010
Proteins with molecular weights of <25 kDa are involved in major biological processes such as ribosome formation, stress adaption (e.g., temperature reduction) and cell cycle control. Despite their importance, the coverage of smaller proteins in standard proteome studies is rather sparse. Here we investigated biochemical and mass spectrometric parameters that influence coverage and validity of identification. The underrepresentation of low molecular weight (LMW) proteins may be attributed to the low numbers of proteolytic peptides formed by tryptic digestion as well as their tendency to be lost in protein separation and concentration/desalting procedures. In a systematic investigation of the LMW proteome of Escherichia coli, a total of 455 LMW proteins (27% of the 1672 listed in the SwissProt protein database) were identified, corresponding to a coverage of 62% of the known cytosolic LMW proteins. Of these proteins, 93 had not yet been functionally classified, and five had not previously been confirmed at the protein level. In this study, the influences of protein extraction (either urea or TFA), proteolytic digestion (solely, and the combined usage of trypsin and AspN as endoproteases) and protein separation (gel- or non-gel-based) were investigated. Compared to the standard procedure based solely on the use of urea lysis buffer, in-gel separation and tryptic digestion, the complementary use of TFA for extraction or endoprotease AspN for proteolysis permits the identification of an extra 72 (32%) and 51 proteins (23%), respectively. Regarding mass spectrometry analysis with an LTQ Orbitrap mass spectrometer, collision-induced fragmentation (CID and HCD) and electron transfer dissociation using the linear ion trap (IT) or the Orbitrap as the analyzer were compared. IT-CID was found to yield the best identification rate, whereas IT-ETD provided almost comparable results in terms of LMW proteome coverage. The high overlap between the proteins identified with IT-CID and IT-ETD allowed the validation of 75% of the identified proteins using this orthogonal fragmentation technique. Furthermore, a new approach to evaluating and improving the completeness of protein databases that utilizes the program RNAcode was introduced and examined.
Journal Article
Assessing the Quality of Cotranscriptional Folding Simulations
by
Kuehnl, Felix
,
Findeiss, Sven
,
Stadler, Peter F
in
Bioinformatics
,
Computer applications
,
Computer programs
2020
Structural changes in RNAs are an important contributor to controlling gene expression not only at the post-transcriptional stage but also during transcription. A subclass of riboswitches and RNA thermometers located in the 5′ region of the primary transcript regulates the downstream functional unit — usually an ORF — through premature termination of transcription. Such elements not only occur naturally but they are also attractive devices in synthetic biology. The possibility to design such riboswitches or RNA thermometers is thus of considerable practical interest. Since these functional RNA elements act already during transcription, it is important to model and understand the dynamics of folding and, in particular, the formation of intermediate structures concurrently with transcription. Cotranscriptional folding simulations are therefore an important step to verify the functionality of design constructs before conducting expensive and labour-intensive wet lab experiments. For RNAs, full-fledged molecular dynamics simulations are far beyond practical reach both because of the size of the molecules and the time scales of interest. Even at the simplified level of secondary structures further approximations are necessary. The BarMap approach is based on representing the secondary structure landscape for each individual transcription step by a coarse-grained representation that only retains a small set of low-energy local minima and the energy barriers between them. The folding dynamics between two transcriptional elongation steps is modeled as a Markov process on this representation. Maps between pairs of consecutive coarse-grained landscapes make it possible to follow the folding process as it changes in response to transcription elongation. In its original implementation, the BarMap software provides a general framework to investigate RNA folding dynamics on temporally changing landscapes. It is, however, difficult to use in particular for specific scenarios such as cotranscriptional folding. To overcome this limitation, we developed the user-friendly BarMap-QA pipeline described in detail in this contribution. It is illustrated here by an elaborate example that emphasizes the careful monitoring of several quality measures. Using an iterative workflow, a reliable and complete kinetics simulation of a synthetic, transcription regulating riboswitch is obtained using minimal computational resources. All programs and scripts used in this contribution are free software and available for download as a source distribution for Linux®, or as a platform-independent Docker® image including support for Apple macOS® and Microsoft Windows®.