bioRxiv ScienceSearch

Biology subjects

Bravo, H. C.

Publications and source records attributed to Bravo, H. C..

2 recordsLinked to original sources

Epiviz File Server: Query, Transform and Interactively Explore Data from Indexed Genomic Files

Genomic data repositories like The Cancer Genome Atlas (TCGA), Encyclopedia of DNA Elements (ENCODE), Bioconductors AnnotationHub and ExperimentHub etc., provide public access to large amounts of genomic data as flat files. Researchers often download a subset of files data from these repositories to perform their data analysis. As these data repositories become larger, researchers often face bottlenecks in their exploratory data analysis. Based on the concepts of a NoDB paradigm, we developed epivizFileServer, a Python library that implements an in-situ data query system for local or remotely hosted indexed genomic files, not only for visualization but also data manipulation. The File Server library decouples data from analysis workflows and provides an abstract interface to define computations independent of the location, format or structure of the file.

bioinformatics

scGAIN: Single Cell RNA-seq Data Imputation using Generative Adversarial Networks

Single cell RNA sequencing (scRNA-seq) provides a rich view into the heterogeneity underlying a cell population. However single-cell data are usually noisy and very sparse due to the presence of dropout genes. In this work we propose an approach to impute missing gene expressions in single cell data using generative adversarial networks (GANs). By learning an approximate distribution of the data, our approach, scGAIN, can impute dropouts in simulated and real single cell data. The work in this paper discusses how to adopt GAIN training model into the domain of imputing single cell data. Experiments show that scGAIN gives competitive results compared to the state-of-the-art approaches while showing superiority in various aspects in simulation and real data. Imputation by scGAIN successfully recovers the underlying clustering of different subpopulations, provides sharp estimates around true mean expressions and increase the correspondence with matched bulk RNAseq experiments.

bioinformatics