Catalogue Search | MBRL
Search Results Heading
Explore the vast range of titles available.
MBRLSearchResults
-
DisciplineDiscipline
-
Is Peer ReviewedIs Peer Reviewed
-
Item TypeItem Type
-
SubjectSubject
-
YearFrom:-To:
-
More FiltersMore FiltersSourceLanguage
Done
Filters
Reset
18
result(s) for
"Samarakoon, Hiruna"
Sort by:
Efficient end-to-end long-read sequence mapping using minimap2-fpga integrated with hardware accelerated chaining
by
Liyanage, Kisaru
,
Gamaarachchi, Hasindu
,
Parameswaran, Sri
in
631/114
,
639/166/987
,
639/705/794
2023
minimap2
is the gold-standard software for reference-based sequence mapping in third-generation long-read sequencing. While
minimap2
is relatively fast, further speedup is desirable, especially when processing a multitude of large datasets. In this work, we present
minimap2-fpga
, a hardware-accelerated version of
minimap2
that speeds up the mapping process by integrating an FPGA kernel optimised for chaining. Integrating the FPGA kernel into
minimap2
posed significant challenges that we solved by accurately predicting the processing time on hardware while considering data transfer overheads, mitigating hardware scheduling overheads in a multi-threaded environment, and optimizing memory management for processing large realistic datasets. We demonstrate speed-ups in end-to-end run-time for data from both Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio).
minimap2-fpga
is up to 79% and 53% faster than
minimap2
for
∼
30
×
ONT and
∼
50
×
PacBio datasets respectively, when mapping without base-level alignment. When mapping with base-level alignment,
minimap2-fpga
is up to 62% and 10% faster than
minimap2
for
∼
30
×
ONT and
∼
50
×
PacBio datasets respectively. The accuracy is near-identical to that of original
minimap2
for both ONT and PacBio data, when mapping both with and without base-level alignment.
minimap2-fpga
is supported on Intel FPGA-based systems (evaluations performed on an on-premise system) and Xilinx FPGA-based systems (evaluations performed on a cloud system). We also provide a well-documented library for the FPGA-accelerated chaining kernel to be used by future researchers developing sequence alignment software with limited hardware background.
Journal Article
GPU accelerated adaptive banded event alignment for rapid comparative nanopore signal analysis
by
Simpson, Jared T.
,
Smith, Martin A.
,
Parameswaran, Sri
in
Adaptive algorithms
,
Algorithms
,
Alignment
2020
Background
Nanopore sequencing enables portable, real-time sequencing applications, including point-of-care diagnostics and in-the-field genotyping. Achieving these outcomes requires efficient bioinformatic algorithms for the analysis of raw nanopore signal data. However, comparing raw nanopore signals to a biological reference sequence is a computationally complex task. The dynamic programming algorithm called Adaptive Banded Event Alignment (ABEA) is a crucial step in polishing sequencing data and identifying non-standard nucleotides, such as measuring DNA methylation. Here, we parallelise and optimise an implementation of the ABEA algorithm (termed
f5c
) to efficiently run on heterogeneous CPU-GPU architectures.
Results
By optimising memory, computations and load balancing between CPU and GPU, we demonstrate how
f5c
can perform ∼3-5 × faster than an optimised version of the original CPU-only implementation of ABEA in the
Nanopolish
software package. We also show that
f5c
enables DNA methylation detection on-the-fly using an embedded System on Chip (SoC) equipped with GPUs.
Conclusions
Our work not only demonstrates that complex genomics analyses can be performed on lightweight computing systems, but also benefits High-Performance Computing (HPC). The associated source code for
f5c
along with GPU optimised ABEA is available at
https://github.com/hasindu2008/f5c
.
Journal Article
Flexible and efficient handling of nanopore sequencing signal data with slow5tools
by
Amos, Timothy G.
,
Deveson, Ira W.
,
Parameswaran, Sri
in
Animal Genetics and Genomics
,
Bioinformatics
,
Biomedical and Life Sciences
2023
Nanopore sequencing is being rapidly adopted in genomics. We recently developed SLOW5, a new file format with advantages for storage and analysis of raw signal data from nanopore experiments. Here we introduce
slow5tools
, an intuitive toolkit for handling nanopore data in SLOW5 format.
Slow5tools
enables lossless data conversion and a range of tools for interacting with SLOW5 files.
Slow5tools
uses multi-threading, multi-processing, and other engineering strategies to achieve fast data conversion and manipulation, including live FAST5-to-SLOW5 conversion during sequencing. We provide examples and benchmarking experiments to illustrate
slow5tools
usage, and describe the engineering principles underpinning its performance.
Journal Article
Chloroplast genome, nuclear ITS regions, mitogenome regions, and Skmer analysis resolved the genetic relationship among Cinnamomum species in Sri Lanka
by
Naranpanawa, Nathasha
,
Chandrasekara, C. H. W. M. R. Bhagya
,
Jayasundara, S.
in
Annotations
,
Bayesian analysis
,
Biology and Life Sciences
2023
Cinnamomum species have gained worldwide attention because of their economic benefits. Among them, C . verum (synonymous with C . zeylanicum Blume), commonly known as Ceylon Cinnamon or True Cinnamon is mainly produced in Sri Lanka. In addition, Sri Lanka is home to seven endemic wild cinnamon species, C . capparu-coronde , C . citriodorum , C . dubium , C . litseifolium , C . ovalifolium , C . rivulorum and C . sinharajaense . Proper identification and genetic characterization are fundamental for the conservation and commercialization of these species. While some species can be identified based on distinct morphological or chemical traits, others cannot be identified easily morphologically or chemically. The DNA barcoding using rbc L, mat K, and trn H -psb A regions could not also resolve the identification of Cinnamomum species in Sri Lanka. Therefore, we generated Illumina Hiseq data of about 20x coverage for each identified species and a C . verum sample (India) and assembled the chloroplast genome, nuclear ITS regions, and several mitochondrial genes, and conducted Skmer analysis. Chloroplast genomes of all eight species were assembled using a seed-based method.According to the Bayesian phylogenomic tree constructed with the complete chloroplast genomes, the C . verum (Sri Lanka) is sister to previously sequenced C . verum (NC_035236.1, KY635878.1), C . dubium and C . rivulorum . The C . verum sample from India is sister to C . litseifolium and C . ovalifolium . According to the ITS regions studied, C . verum (Sri Lanka) is sister to C . verum (NC_035236.1), C . dubium and C . rivulorum . Cinnamomum verum (India) shares an identical ITS region with C . ovalifolium , C . litseifolium , C . citriodorum , and C . capparu-coronde . According to the Skmer analysis C . verum (Sri Lanka) is sister to C . dubium and C . rivulorum , whereas C. verum (India) is sister to C . ovalifolium , and C . litseifolium . The chloroplast gene ycf1 was identified as a chloroplast barcode for the identification of Cinnamomum species. We identified an 18 bp indel region in the ycf1 gene, that could differentiate C . verum (India) and C . verum (Sri Lanka) samples tested.
Journal Article
Fast nanopore sequencing data analysis with SLOW5
by
Amos, Timothy G.
,
Saadat, Hassaan
,
Smith, Martin A.
in
631/61/514/1948
,
692/308/2056
,
Agriculture
2022
Nanopore sequencing depends on the FAST5 file format, which does not allow efficient parallel analysis. Here we introduce SLOW5, an alternative format engineered for efficient parallelization and acceleration of nanopore data analysis. Using the example of DNA methylation profiling of a human genome, analysis runtime is reduced from more than two weeks to approximately 10.5 h on a typical high-performance computer. SLOW5 is approximately 25% smaller than FAST5 and delivers consistent improvements on different computer architectures.
Nanopore sequencing data are rapidly analyzed with parallel data access.
Journal Article
Genopo: a nanopore sequencing analysis toolkit for portable Android devices
by
Punchihewa, Sanoj
,
Deveson, Ira W.
,
Hammond, Jillian M.
in
45/23
,
631/114/2398
,
631/61/514/1948
2020
The advent of portable nanopore sequencing devices has enabled DNA and RNA sequencing to be performed in the field or the clinic. However, advances in in situ genomics require parallel development of portable, offline solutions for the computational analysis of sequencing data. Here we introduce
Genopo
, a mobile toolkit for nanopore sequencing analysis.
Genopo
compacts popular bioinformatics tools to an Android application, enabling fully portable computation. To demonstrate its utility for in situ genome analysis, we use
Genopo
to determine the complete genome sequence of the human coronavirus SARS-CoV-2 in nine patient isolates sequenced on a nanopore device, with
Genopo
executing this workflow in less than 30 min per sample on a range of popular smartphones. We further show how
Genopo
can be used to profile DNA methylation in a human genome sample, illustrating a flexible, efficient architecture that is suitable to run many popular bioinformatics tools and accommodate small or large genomes. As the first ever smartphone application for nanopore sequencing analysis,
Genopo
enables the genomics community to harness this cheap, ubiquitous computational resource.
Samarakoon et al. present
Genopo
, a portable toolkit for analysis of nanopore sequencing data on Android smartphone and tablet devices.
Genopo
enables field-based nanopore analysis using widely accessible technology.
Journal Article
Optimised Computational and Visualisation Methods to Advance Nanopore Raw Signal Analysis
DNA sequencing has transformed medicine by allowing researchers to identify genetic problems, create tailored treatment plans, and develop targeted drugs. There are different sequencing technologies such as Sanger, Illumina, PacBio, and Nanopore, each with its own strengths for different genetic studies. These technologies produce massive amounts of data, which brings both challenges and opportunities for storing, analysing, and understanding the genome.Among other sequencing platforms, Oxford Nanopore Technologies’s (ONT) nanopore sequencing emerges as a disruptive sequencing method with the unique ability to sequence both short and long native DNA and RNA. Nanopore sequencers generate electrical raw signals as DNA or RNA molecules pass through nanopores, which are then decoded (basecalled) to determine the nucleotide (A,C,G,T/U) sequence. These nanopore raw signals are also useful for downstream analyses such as studying DNA methylation, DNA damage, and RNA secondary structures. However, the computational and visualisation methods for analysing these raw signal data are immature, presenting challenges in the utility and accessibility of this valuable data.The hundreds of terabytes of raw nanopore data generated by ONT sequencing undergo a complex lifecycle, transforming to suit various analysis requirements. This necessitates effective management strategies to ensure smooth data handling throughout the process. This thesis introduces optimised methods to address these challenges, including data conversion, merging, splitting, and dataset summarising. These techniques are designed to be scalable and adaptable to various computational platforms, ensuring efficient and flexible data management for nanopore sequencing.Basecalling, the process of converting raw nanopore signals to nucleotide sequences, is a computationally intensive process often performed on GPUs. As the initial step in most ONT analysis pipelines, optimising basecalling is crucial for efficient use of computational resources. This thesis addresses an input/output (I/O) bottleneck identified within the basecalling process that limited its scalability.Similar to read-to-reference alignment, signal-to-sequence alignment is a fundamental operation in the nanopore domain. Many existing alignment methods operate at the k-mer level, limiting their resolution for tasks such as base modification detection. Re-aligningthese to the most informative base within the k-mer is crucial. Additionally, some alignments, like basecaller move tables, are limited to signal-to-read scope and require transformation to signal-to-reference. This theses introduces novel algorithms to address above challenges and presents a comprehensive framework equipped with fast data retrieval methods for interactive signal-to-sequence alignment visualisation.Most nanopore signal aligners utilise a k-mer table that indexes k-mers and their corresponding average raw current levels. Custom k-mer tables are essential for accurate signal alignment and interpretation, particularly when default models are unavailable or inadequate for specific sequencing conditions. Therefore, this thesis also presents a de novo k-mer table modelling method to deduce a lightweight k-mer table that yields comparable results to larger, default models.In summary this thesis presents optimised computational and visualisation methods to advance nanopore raw signal analysis. These open-source efforts adhere to bioinformatics software best practices, including version control, documentation, testing and validation. Moreover, the presented work is refined for performance on standard workstation computers, high performance servers and cloud computing platforms.
Dissertation
The enduring advantages of the SLOW5 file format for raw nanopore sequencing data
2025
Nanopore sequencing is a widespread and important method in genomics science. The raw electrical current signal data from a typical nanopore sequencing experiment is large and complex. This can be stored in two alternative file formats that are presently supported: POD5 is a signal data file format used by default on instruments from Oxford Nanopore Technologies (ONT); SLOW5 is an open-source file format originally developed as an alternative to ONT's previous file format, which was known as FAST5. The choice of format may have important implications for the cost, speed and simplicity of nanopore signal data analysis, management and storage. To inform this choice, we present a comparative evaluation of POD5 vs SLOW5. We conducted benchmarking experiments assessing file size, analysis performance and usability on a variety of different computer architectures. SLOW5 showed superior performance during sequential and non-sequential (random access) file reading on most systems, manifesting in faster, cheaper basecalling and other analysis, and we could find no instance in which POD5 file reading was significantly faster than SLOW5. We demonstrate that SLOW5 file writing is highly parallelisable, thereby meeting the demands of data acquisition on ONT instruments. Our analysis also identified differences in the complexity and stability of the software libraries for SLOW5 (slow5lib) and POD5 (pod5), including a large discrepancy in the number of underlying software dependencies, which may complicate the pod5 compilation process. In summary, many of the advantages originally conceived for SLOW5 remain relevant today, despite the replacement of FAST5 with POD5 as ONT's core file format.
minimap2-fpga: Integrating hardware-accelerated chaining for efficient end-to-end long-read sequence mapping
2023
minimap2 is the gold-standard software for reference-based sequence mapping in third-generation long-read sequencing. While minimap2 is relatively fast, further speedup is desirable, especially when processing a multitude of large datasets. In this work, we present minimap2-fpga, a hardware-accelerated version of minimap2 that speeds up the mapping process by integrating an FPGA kernel optimised for chaining. We demonstrate speed-ups in end-to-end run-time for data from both Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio). minimap2-fpga is up to 79% and 53% faster than minimap2 for ∼ 30× ONT and ∼ 50× PacBio datasets respectively, when mapping without base-level alignment. When mapping with base-level alignment, minimap2-fpga is up to 62% and 10% faster than minimap2 for ∼ 30× ONT and ∼ 50× PacBio datasets respectively. The accuracy is near-identical to that of original minimap2 for both ONT and PacBio data, when mapping both with and without base-level alignment. minimap2-fpga is supported on Intel FPGA-based systems (evaluations performed on an on-premise system) and Xilinx FPGA-based systems (evaluations performed on a cloud system). We also provide a well-documented library for the FPGA-accelerated chaining kernel to be used by future researchers developing sequence alignment software with limited hardware background.