bioRxiv ScienceSearch

Biology subjects

Chen, S.-H.

Publications and source records attributed to Chen, S.-H..

3 recordsLinked to original sources

Visualisation and analysis of RNA-Seq assembly graphs

RNA-sequencing (RNA-Seq) is a powerful transcriptome profiling technology enabling transcript discovery and quantification. RNA-Seq data are large, and most commonly used as a source of genelevel quantification measurements, whilst the underlying assemblies of reads, if inspected, are usually viewed as sequence reads mapped on to a reference genome. Whilst sufficient for many needs, when the underlying transcript assemblies are complex, this visualisation approach can be limiting; errors in assembly can be difficult to spot and interpretation of splicing events is challenging.\n\nHere we report on the development of a graph-based visualisation method as a complementary approach to understanding transcript diversity and read assembly from short-read RNA-Seq data. Following the mapping of reads to the reference genome, read-to-read comparison is performed on all reads mapping to a given gene, producing a matrix of weighted similarity scores between reads. This is used to produce an RNA assembly graph where nodes represent reads derived from a cDNA and edges similarity scores between reads, above a defined threshold. Visualisation of resulting graphs is performed using Graphia Professional. This tool can render the often large and complex graph topologies that result from DNA/RNA sequence assembly in 3D space and supports info rmatio no verlay on to nodes, e.g. transcript models. We have also implemented an analysis pipeline for the creation of RNA assembly graphs with both a command-line and web-based interface that allows users to create and visualise these data. Here we demonstrate the utility of this approach on RNA-Seq data, including the unusual structure of these graphs and how they can be used to identify issues in assembly, repetitive sequences within transcripts and splice variants. We believe this approach has the potential to significantly improve our understanding of transcript complexity.

bioinformatics

Progression of chronic kidney disease in African American with type 2 diabetes mellitus using topology learning in electronic medical records

BackgroundChronic kidney disease (CKD) is a common, complex, and heterogeneous disease impacting aging populations. Determining the landscape of disease progression trajectories from midlife to senior age in a \"real-world\" context allows us to better understand the progression of CKD, the heterogeneity of progression patterns among the risk population, and the interactions with other clinical conditions. Genetics also plays an important role. In previous work, we and others have demonstrated that African Americans with high-risk APOL1 genotypes are more likely to develop CKD, tend to develop CKD earlier, and the disease progresses faster. Diabetes, which is more common in African Americans, also significantly increases risk for CKD.\n\nData and MethodElectronic medical records (EMRs) were used to outline the first CKD progression trajectory roadmap for an African American population with type 2 diabetes. By linking participants in 5 genome-wide association study (GWAS) to their clinical records at Wake Forest Baptist Medical Center (WFBMC), an EMR-GWAS cohort was established (n = 1,581). Patients health status was described by 18 Essential Clinical Indices across 84,009 clinical encounters. A novel graph learning algorithm, Discriminative Dimensionality Reduction Tree (DDRTree) was implemented, to establish the trajectories of declines in health. Moreover, a prediction model for new patients was proposed along the learned graph structure. We annotated these trajectories with clinical and genomic features including kidney function, other major risk indices of CKD, APOL1 genotypes, and age. The prediction power of the learned disease progression trajectories was further examined using the k-nearest neighbor model.\n\nResultsThe CKD progression trajectory roadmap revealed diverse kidney failure pathways associated with different clinical conditions. Specifically, we identified one high-risk trajectory and two low-risk trajectories. Switching pathways from low-risk trajectories to the high-risk one was associated with accelerated decline in kidney function. On this roadmap, patients with APOL1 high-risk genotypes were enriched in the high-risk trajectory, suggesting fundamentally different disease progression mechanisms from those without APOL1 risk genotypes. The k-nearest neighbor-based prediction showed effective prediction rate of 87%.\n\nConclusionThe CKD progression trajectory roadmap revealed novel diverse renal failure pathways in African Americans with type 2 diabetes mellitus and highlights disease progression patterns that associate with APOL1 renal-risk genotypes.

bioinformatics

Assembly of a Parts List of the Human Mitotic Cell Cycle Machinery

The set of proteins required for mitotic division remains poorly characterised. Here, an extensive series of correlation analyses of human and mouse transcriptomics data was performed to identify genes strongly and reproducibly associated with cells undergoing S/G2-M phases of the cell cycle. In so doing, a list of 701 cell cycle-associated genes was defined and shown that whilst many are only expressed during these phases, the expression of others is also driven by alternative promoters. Of this list, 496 genes have known cell cycle functions, whereas 205 were assigned as putative cell cycle genes, 53 of which are functionally uncharacterised. Among these, 27 were screened for subcellular localisation revealing many to be nuclear localised and at least four to be novel centrosomal proteins. Furthermore, 10 others inhibited cell proliferation upon siRNA knockdown. This study presents the first comprehensive list of human cell cycle proteins, identifying many new candidate proteins.

genomics