bioRxiv ScienceSearch

Biology subjects

Pers, T. H.

Publications and source records attributed to Pers, T. H..

3 recordsLinked to original sources

scVAE: Variational auto-encoders for single-cell gene expression data

MotivationModels for analysing and making relevant biological inferences from massive amounts of complex single-cell transcriptomic data typically require several individual data-processing steps, each with their own set of hyperparameter choices. With deep generative models one can work directly with count data, make likelihood-based model comparison, learn a latent representation of the cells and capture more of the variability in different cell populations.\n\nResultsWe propose a novel method based on variational auto-encoders (VAEs) for analysis of single-cell RNA sequencing (scRNA-seq) data. It avoids data preprocessing by using raw count data as input and can robustly estimate the expected gene expression levels and a latent representation for each cell. We tested several count likelihood functions and a variant of the VAE that has a priori clustering in the latent space. We show for several scRNA-seq data sets that our method outperforms recently proposed scRNA-seq methods in clustering cells and that the resulting clusters reflect cell types.\n\nAvailability and implementationOur method, called scVAE, is implemented in Python using the TensorFlow machine-learning library, and it is freely available at https://github.com/scvae/scvae.

bioinformatics

PAIRUP-MS: Pathway Analysis and Imputation to Relate Unknowns in Profiles from Mass Spectrometry-based metabolite data

Metabolomics is a powerful approach for discovering biomarkers and metabolic quantitative trait loci. While untargeted profiling methods can measure up to thousands of metabolite signals in a single experiment, many signals cannot be readily identified as known metabolites or compared across datasets, making it difficult to infer biology and to conduct well-powered meta-analyses across studies. To deal with these challenges, we developed a suite of computational methods, PAIRUP-MS, to match metabolite signals across mass spectrometry-based profiling datasets using an imputation-based approach and to generate pathway annotations for these signals. We performed meta and pathway analyses for both known and unknown signals in multiple datasets and then validated the results using genetic associations. Finally, we applied the methods to detect metabolite signals and pathways associated with body mass index, demonstrating that our framework is useful for analyzing unknown signals in a robust and biologically meaningful manner and for improving the power of untargeted metabolomics studies.

bioinformatics

A Comprehensive Reanalysis Of Publicly Available GWAS Datasets Reveals An X Chromosome Rare Regulatory Variant Associated With High Risk For Type 2 Diabetes.

The reanalysis of publicly available GWAS data represents a powerful and cost-effective opportunity to gain insights into the genetics and pathophysiology of complex diseases. We demonstrate this by gathering and reanalyzing public type 2 diabetes (T2D) GWAS data for 70,127 subjects, using an innovative imputation and association strategy based on multiple reference panels (1000G and UK10K). This approach led us replicate and fine map 50 known T2D loci, and identify seven novel associated regions: five driven by common variants in or near LYPLAL1, NEUROG3, CAMKK2, ABO and GIP genes; one by a low frequency variant near EHMT2; and one driven by a rare variant in chromosome Xq23, associated with a 2.7-fold increased risk for T2D in males, and located within an active enhancer associated with the expression of Angiotensin II Receptor type 2 gene (AGTR2), a known modulator of insulin sensitivity. We further show that the risk T allele reduces binding of a nuclear protein, resulting in increased enhancer activity in muscle cells. Beyond providing novel insights into the genetics and pathophysiology of T2D, these results also underscore the value of reanalyzing publicly available data using novel analytical approaches.

genetics