Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
975
result(s) for
"random sequencing"
Sort by:
Next-Generation Sequencing for the Detection of Microbial Agents in Avian Clinical Samples
2023
Direct-targeted next-generation sequencing (tNGS), with its undoubtedly superior diagnostic capacity over real-time PCR (RT-PCR), and direct-non-targeted NGS (ntNGS), with its higher capacity to identify and characterize multiple agents, are both likely to become diagnostic methods of choice in the future. tNGS is a rapid and sensitive method for precise characterization of suspected agents. ntNGS, also known as agnostic diagnosis, does not require a hypothesis and has been used to identify unsuspected infections in clinical samples. Implemented in the form of multiplexed total DNA metagenomics or as total RNA sequencing, the approach produces comprehensive and actionable reports that allow semi-quantitative identification of most of the agents present in respiratory, cloacal, and tissue samples. The diagnostic benefits of the use of direct tNGS and ntNGS are high specificity, compatibility with different types of clinical samples (fresh, frozen, FTA cards, and paraffin-embedded), production of nearly complete infection profiles (viruses, bacteria, fungus, and parasites), production of “semi-quantitative” information, direct agent genotyping, and infectious agent mutational information. The achievements of NGS in terms of diagnosing poultry problems are described here, along with future applications. Multiplexing, development of standard operating procedures, robotics, sequencing kits, automated bioinformatics, cloud computing, and artificial intelligence (AI) are disciplines converging toward the use of this technology for active surveillance in poultry farms. Other advances in human and veterinary NGS sequencing are likely to be adaptable to avian species in the future.
Journal Article
Scheduling with Opting Out: Improving upon Random Priority
2001
In a scheduling problem where agents can opt out, we show that the familiar random priority (RP) mechanism can be improved upon by another mechanism dubbed probabilistic serial (PS). Both mechanisms are nonmanipulable in a strong sense, but the latter is Pareto superior to the former and serves a larger (expected) number of agents. The PS equilibrium outcome is easier to compute than the RP outcome; on the other hand, RP is easier to implement than PS. We show that the improvement of PS over RP is significant but small: at most a couple of percentage points in the relative welfare gain and the relative difference in quantity served. Both gains vanish when the number of agents is large; hence both mechanisms can be used as a proxy of each other.
Journal Article
Changes in the levels of mRNAs for putative cell growth-related genes in the albedo and flavedo during citrus fruit development
2000
Changes in mRNA levels for the seven gene homologues to endoxyloglucan transferase-related protein, expansin, extensin, β-1,3 glucanase, glycine-rich protein, pectinacetylesterase and pectinesterase, which were obtained by random sequencing studies, were investigated in relation to rind development in citrus (Citrus unshiu Marc.) fruit. Expression patterns in the albedo and flavedo were classified into four types: Type-I, transcript levels low in early fruit development but increased at the ripening stage; Type-II, transcript levels high until mid-development and then decreased towards ripening; Type-III, detectable transcript limited to fruitlets at 26 days after flowering (DAF); Type-IV, ubiquitous during the development. Based on the expression patterns, we discuss the possible roles of these genes in rind development.
Journal Article
Construction and Sequence Analysis of Subtraction Complementary DNA Libraries from Human Preimplantation Embryos
by
Gindilis, Viktor
,
Strom, Charles
,
Morozov, Gennady
in
Adenosine triphosphatase
,
Algorithms
,
Antigens
1999
Because stage-specific genetic expression in human preimplantation development is not sufficiently studied, we have undertaken the construction of a subtraction complementary DNA (cDNA) library enriched for transcripts specific for human blastocysts.
For this purpose individual pools of cDNAs synthesized from four hatched blastocysts and three cleaving 8- to 10-cell embryos were exposed to suppression subtractive hybridization to minimize the presence of transcripts of housekeeping genes and other genes of maternal origin known to be expressed earlier in preimplantation development. Random clones of this library were sequenced and analyzed using the BLAST algorithm.
The resulting subtraction library had a complexity of 3 x 10(5) and an average size of inserts of about 0.8 kb. Sequencing of random library clones revealed the following human genes: CD9 antigen, fatty acid binding protein, ferritin heavy chain, amyloid precursor, MAP kinase messenger RNAs, DNA clone 127H14, messenger RNA for diacylglycerol kinase, a sequence homologous to C1 inhibitor, messenger RNA for the KIAA0145 gene, and others.
The presence of these genes in human preimplantation development suggests expression specific to the blastocyst stage.
Journal Article
Determination of high-confidence germline genetic variants in next-generation sequencing through machine learning models: an approach to reduce the burden of orthogonal confirmation
2025
Background
Orthogonal confirmation of variants identified by next-generation sequencing (NGS) is routinely performed in many clinical laboratories to improve assay specificity. However, confirmatory testing of all clinically significant variants increases both turnaround time and operating costs for laboratories. Improvements to early NGS methods and bioinformatics algorithms have dramatically improved variant calling accuracy, particularly for single nucleotide variants (SNVs), thus calling into question the necessity of confirmatory testing for all variant types. The purpose of this study is to develop a new machine learning approach to capture false positive heterozygous variants (SNVs) from whole exome sequencing (WES) data.
Results
WES variant calls from Genome in a Bottle (GIAB) cell lines and their associated quality features were used to train five different machine learning models to predict whether a variant was a true positive or false positive based on quality metrics. Logistic regression and random forest models exhibited the highest false positive capture rates among the selected models, but GradientBoosting achieved the best balance between false positive capture rates and true positive flag rates. Further assessment using simulated false positive events as well as different combinations of quality features showed that model performance can be refined. Integration of the highest-performing models into a custom two-tiered confirmation bypass pipeline with additional guardrail metrics achieved 99.9% precision and 98% specificity in the identification of true positive heterozygous SNVs within the GIAB benchmark regions. Furthermore, testing on an independent set of heterozygous SNVs (
n
= 93) detected by exome sequencing of patient samples and cell lines demonstrated 100% accuracy.
Conclusions
Machine-learning models can be trained to classify SNVs into high or low-confidence categories with high precision, thus reducing the level of confirmatory testing required. Laboratories interested in deploying such models should consider incorporating additional quality criteria and thresholds to serve as guardrails in the assessment process.
Journal Article
A general and flexible method for signal extraction from single-cell RNA-seq data
2018
Single-cell RNA-sequencing (scRNA-seq) is a powerful high-throughput technique that enables researchers to measure genome-wide transcription levels at the resolution of single cells. Because of the low amount of RNA present in a single cell, some genes may fail to be detected even though they are expressed; these genes are usually referred to as dropouts. Here, we present a general and flexible zero-inflated negative binomial model (ZINB-WaVE), which leads to low-dimensional representations of the data that account for zero inflation (dropouts), over-dispersion, and the count nature of the data. We demonstrate, with simulated and real data, that the model and its associated estimation procedure are able to give a more stable and accurate low-dimensional representation of the data than principal component analysis (PCA) and zero-inflated factor analysis (ZIFA), without the need for a preliminary normalization step.
Single-cell RNA sequencing (scRNA-seq) data provides information on transcriptomic heterogeneity within cell populations. Here, Risso et al develop ZINB-WaVE for low-dimensional representations of scRNA-seq data that account for zero inflation, over-dispersion, and the count nature of the data.
Journal Article
Taxanorm: a novel taxa-specific normalization approach for microbiome data
2024
Background
In high-throughput sequencing studies, sequencing depth, which quantifies the total number of reads, varies across samples. Unequal sequencing depth can obscure true biological signals of interest and prevent direct comparisons between samples. To remove variability due to differential sequencing depth, taxa counts are usually normalized before downstream analysis. However, most existing normalization methods scale counts using size factors that are sample specific but not taxa specific, which can result in over- or under-correction for some taxa.
Results
We developed TaxaNorm, a novel normalization method based on a zero-inflated negative binomial model. This method assumes the effects of sequencing depth on mean and dispersion vary across taxa. Incorporating the zero-inflation part can better capture the nature of microbiome data. We also propose two corresponding diagnosis tests on the varying sequencing depth effect for validation. We find that TaxaNorm achieves comparable performance to existing methods in most simulation scenarios in downstream analysis and reaches a higher power for some cases. Specifically, it balances power and false discovery control well. When applying the method in a real dataset, TaxaNorm has improved performance when correcting technical bias.
Conclusion
TaxaNorm both sample- and taxon- specific bias by introducing an appropriate regression framework in the microbiome data, which aids in data interpretation and visualization. The ‘TaxaNorm’ R package is freely available through the CRAN repository
https://CRAN.R-project.org/package=TaxaNorm
and the source code can be downloaded at
https://github.com/wangziyue57/TaxaNorm
.
Journal Article
Random access in large-scale DNA data storage
2018
200 MB of digital data is stored in DNA, randomly accessed and recovered using an error-free approach.
Synthetic DNA is durable and can encode digital data with high density, making it an attractive medium for data storage. However, recovering stored data on a large-scale currently requires all the DNA in a pool to be sequenced, even if only a subset of the information needs to be extracted. Here, we encode and store 35 distinct files (over 200 MB of data), in more than 13 million DNA oligonucleotides, and show that we can recover each file individually and with no errors, using a random access approach. We design and validate a large library of primers that enable individual recovery of all files stored within the DNA. We also develop an algorithm that greatly reduces the sequencing read coverage required for error-free decoding by maximizing information from all sequence reads. These advances demonstrate a viable, large-scale system for DNA data storage and retrieval.
Journal Article
U1 snRNP regulates chromatin retention of noncoding RNAs
2020
Long noncoding RNAs (lncRNAs) and promoter- or enhancer-associated unstable transcripts locate preferentially to chromatin, where some regulate chromatin structure, transcription and RNA processing
1
–
13
. Although several RNA sequences responsible for nuclear localization have been identified—such as repeats in the lncRNA
Xist
and Alu-like elements in long RNAs
14
–
16
—how lncRNAs as a class are enriched at chromatin remains unknown. Here we describe a random, mutagenesis-coupled, high-throughput method that we name ‘RNA elements for subcellular localization by sequencing’ (mutREL-seq). Using this method, we discovered an RNA motif that recognizes the U1 small nuclear ribonucleoprotein (snRNP) and is essential for the localization of reporter RNAs to chromatin. Across the genome, chromatin-bound lncRNAs are enriched with 5′ splice sites and depleted of 3′ splice sites, and exhibit high levels of U1 snRNA binding compared with cytoplasm-localized messenger RNAs. Acute depletion of U1 snRNA or of the U1 snRNP protein component SNRNP70 markedly reduces the chromatin association of hundreds of lncRNAs and unstable transcripts, without altering the overall transcription rate in cells. In addition, rapid degradation of SNRNP70 reduces the localization of both nascent and polyadenylated lncRNA transcripts to chromatin, and disrupts the nuclear and genome-wide localization of the lncRNA
Malat1
. Moreover, U1 snRNP interacts with transcriptionally engaged RNA polymerase II. These results show that U1 snRNP acts widely to tether and mobilize lncRNAs to chromatin in a transcription-dependent manner. Our findings have uncovered a previously unknown role of U1 snRNP beyond the processing of precursor mRNA, and provide molecular insight into how lncRNAs are recruited to regulatory sites to carry out chromatin-associated functions.
Long noncoding RNAs and certain unstable transcripts tend to localize to chromatin, in a process that is shown here to depend on an RNA motif that recognizes the small nuclear ribonuclear protein U1, and to rely on transcription.
Journal Article
CMIC: an efficient quality score compressor with random access functionality
2022
Background
Over the past few decades, the emergence and maturation of new technologies have substantially reduced the cost of genome sequencing. As a result, the amount of genomic data that needs to be stored and transmitted has grown exponentially. For the standard sequencing data format, FASTQ, compression of the quality score is a key and difficult aspect of FASTQ file compression. Throughout the literature, we found that the majority of the current quality score compression methods do not support random access. Based on the above consideration, it is reasonable to investigate a lossless quality score compressor with a high compression rate, a fast compression and decompression speed, and support for random access.
Results
In this paper, we propose CMIC, an adaptive and random access supported compressor for lossless compression of quality score sequences. CMIC is an acronym of the four steps (classification, mapping, indexing and compression) in the paper. Its framework consists of the following four parts: classification, mapping, indexing, and compression. The experimental results show that our compressor has good performance in terms of compression rates on all the tested datasets. The file sizes are reduced by up to 21.91% when compared with LCQS. In terms of compression speed, CMIC is better than all other compressors on most of the tested cases. In terms of random access speed, the CMIC is faster than the LCQS, which provides a random access function for compressed quality scores.
Conclusions
CMIC is a compressor that is especially designed for quality score sequences, which has good performance in terms of compression rate, compression speed, decompression speed, and random access speed. The CMIC can be obtained in the following way:
https://github.com/Humonex/Cmic
.
Journal Article