bioRxiv ScienceSearch

Biology subjects

Ho, J. W. K.

Publications and source records attributed to Ho, J. W. K..

6 recordsLinked to original sources

Ultrafast clustering of single-cell flow cytometry data using FlowGrid

Flow cytometry is a popular technology for quantitative single-cell profiling of cell surface markers. It enables expression measurement of tens of cell surface protein markers in millions of single cells. It is a powerful tool for discovering cell sub-populations and quantifying cell population heterogeneity. Traditionally, scientists use manual gating to identify cell types, but the process is subjective and is not effective for large multidimensional data. Many clustering algorithms have been developed to analyse these data but most of them are not scalable to very large data sets with more than ten million cells.\n\nHere, we present a new clustering algorithm that combines the advantages of density-based clustering algorithm DBSCAN with the scalability of grid-based clustering. This new clustering algorithm is implemented in python as an open source package, FlowGrid. FlowGrid is memory efficient and scales linearly with respect to the number of cells. We have evaluated the performance of FlowGrid against other state-of-the-art clustering programs and found that FlowGrid produces similar clustering results but with substantially less time. For example, FlowGrid is able to complete a clustering task on a data set of 23.6 million cells in less than 12 seconds, while other algorithms take more than 500 seconds or get into error.\n\nFlowGrid is an ultrafast clustering algorithm for large single-cell flow cy-tometry data. The source code is available at https://github.com/VCCRI/FlowGrid.

bioinformatics

Scavenger: A pipeline for recovery of unaligned reads utilising similarity with aligned reads

MotivationRead alignment is an important step in RNA-seq analysis as the result of alignment forms the basis for further downstream analyses. However, recent studies have shown that published alignment tools have variable mapping sensitivity and do not necessarily align reads which should have been aligned, a problem we termed as the false-negative non-alignment problem.\n\nResultsWe have developed Scavenger, a pipeline for recovering unaligned reads using a novel mechanism which utilises information from aligned reads. Scavenger performs recovery of unaligned reads by re-aligning unaligned reads against a putative location derived from aligned reads with sequence similarity against unaligned reads. We show that Scavenger can successfully recover unaligned reads in both simulated and real RNA-seq datasets, including single-cell RNA-seq data. The reads recovered contain more genetic variants compared to previously aligned reads, indicating that divergence between personal and reference genomes plays a role in the false-negative non-alignment problem. We also explored the impact of read recovery on downstream analyses, in particular gene expression analysis, and showed that Scavenger is able to both recover genes which were previously non-expressed and also increase gene expression, with lowly expressed genes having the most impact from the addition of recovered reads. We also found that the majority of genes with >1 fold change in expression after recovery are categorised as pseudogenes, indicating that pseudogene expression can be affected by the false-negative non-alignment problem. Scavenger helps to solve the false-negative non-alignment problem through recovery of unaligned reads using information from previously aligned reads.\n\nAvailabilityScavenger is available via an open source license in https://github.com/VCCRI/Scavenger/\n\nContactj.ho@victorchang.edu.au

bioinformatics

CardiacProfileR: An R package for extraction and visualisation of heart rate profiles from wearable fitness trackers

A persons heart rate profile, which consists of resting heart rate, increase of heart rate upon exercise and recovery of heart rate after exercise, is traditionally measured by electrocardiography during a controlled exercise stress test. A heart rate profile is a useful clinical tool to identify individuals at risk of sudden death and other cardiovascular conditions. Nonetheless, conducting such exercise stress tests routinely is often inconvenient and logistically challenging for patients. The widespread availability of affordable wearable fitness trackers, such as Fitbit and Apple Watch, provides an exciting new means to collect longitudinal heart rate and physical activity data. We reason that by combining the heart rate and physical activity data from these devices, we can construct a persons heart rate profile. Here we present an open source R package CardiacProfileR for extraction, analysis and visualisation of heart rate dynamics during physical activities from data generated from common wearable heart rate monitors. This package represents a step towards quantitative deep phenotyping in humans. CardiacProfileR is available via an MIT license at https://github.com/VCCRI/CardiacProfileR.

bioinformatics

Ularcirc: Visualisation and enhanced analysis of circular RNAs via back and canonical forward splicing

Circular RNAs (circRNAs) are a unique class of transcripts that can only be identified from sequence alignments spanning discordant junctions, commonly referred to as backsplice junctions (BSJ). The challenges of detecting a BSJ from short read high throughput sequencing (HTS) data has steered software development to focus primarily on algorithmic methods to accurately capture BSJs. Here we present Ularcirc, the first software tool that provides a complete circRNA workflow from detection, integrated visualization, quality filtering of BSJ and forward splicing junctions (FSJ), through to sequence retrieval and downstream functional analysis. More importantly, Ularcirc uses an innovative method to filter out false positive circRNAs coined read alignment distribution (RAD) score which allows detection of circRNAs independent of gene annotations. We used Ularcirc to characterise circRNAs from public and in-house generated data sets and demonstrate how to discover (i) novel splicing patterns of parental transcripts, (ii) internal splicing patterns of circRNA, and (iii) the complexity of BSJ formation. Furthermore, we identify circRNAs that have potential open reading frames longer than their linear sequence. Finally, we have identified and validated the presence of a novel class of circRNA generated from ApoA4 transcripts whose BSJ derive from multiple sites within coding exons. Ularcirc can be accessed via https://github.com/VCCRI/Ularcirc.

bioinformatics

GEOracle: Mining perturbation experiments using free text metadata in Gene Expression Omnibus

There exists over 2.5 million publicly available gene expression samples across 101,000 data series in NCBIs Gene Expression Omnibus (GEO) database. Due to the lack of the use of standardised ontology terms in GEOs free text metadata to annotate the experimental type and sample type, this database remains di[ffi]cult to harness computationally without significant manual intervention.\n\nIn this work, we present an interactive R/Shiny tool called GEOracle that utilises text mining and machine learning techniques to automatically identify perturbation experiments, group treatment and control samples and perform differential expression. We present applications of GEOracle to discover conserved signalling pathway target genes and identify an organ specific gene regulatory network.\n\nGEOracle is effective in discovering perturbation gene targets in GEO by harnessing its free text metadata. Its effectiveness and applicability has been demonstrated by cross validation and two real-life case studies. It opens up new avenues to unlock the gene regulatory information embedded inside large biological databases such as GEO. GEOracle is available at https://github.com/VCCRI/GEOracle.

bioinformatics

Impact Of Sequencing Depth And Read Length On Single Cell RNA Sequencing Data: Lessons From T Cells

Single cell RNA sequencing (scRNA-seq) has shown great potential in measuring the gene expression profiles of heterogeneous cell populations. In immunology, scRNA-seq allowed the characterisation of transcript sequence diversity of functionally relevant sub-populations of T cells, and notably the identification of the full length T cell receptor (TCR{beta}), which defines the specificity against cognate antigens. Several factors, such as RNA library capture, cell quality, and sequencing output have been suggested to affect the quality of scRNA-seq data, but these factors have not been systematically examined.\n\nWe studied the effect of read length and sequencing depth on the quality of gene expression profiles, cell type identification, and TCR{beta} reconstruction, utilising 1,305 publically available scRNA-seq datasets, and simulation-based analyses. Gene expression was characterised by an increased number of unique genes identified with short read lengths (<50 bp), but these featured higher technical variability compared to profiles from longer reads. TCR{beta} were detected in 1,027 cells (79%), with a success rate between 81% and 100% for datasets with at least 250,000 (PE) reads of length >50 bp.\n\nSufficient read length and sequencing depth can control technical noise to enable accurate identification of TCR{beta} and gene expression profiles from scRNA-seq data of T cells.

bioinformatics