bioRxiv ScienceSearch

Biology subjects

Yudi Pawitan

Publications and source records attributed to Yudi Pawitan.

3 recordsLinked to original sources

Isoform-level gene expression patterns in single-cell RNA-sequencing data

RNA-sequencing of single-cells enables characterization of transcriptional heterogeneity in seemingly homogenous cell populations. In this study we propose and apply a novel method, ISOform-Patterns (ISOP), based on mixture modeling, to characterize the expression patterns of pairs of isoforms from the same gene in single-cell isoform-level expression data. We define six principal patterns of isoform expression relationships and introduce the concept of differential pattern analysis. We applied ISOP for analysis of single-cell RNA-sequencing data from a breast cancer cell line, with replication in two independent datasets. In the primary dataset we detected and assigned pattern type of 16562 isoform-pairs from 4929 genes. Our results showed that 78% of the isoform pairs displayed a mutually exclusive expression pattern, 14% of the isoform pairs displayed bimodal isoform preference and 8% isoform pairs displayed isoform preference. 26% of the isoform-pair patterns were significant, while remaining isoform-pair patterns can be understood as effects of transcriptional bursting, drop-out and biological heterogeneity. 32% of genes discovered through differential pattern analysis were novel and not detected by differential expression analysis. ISOP provides a novel approach for characterization of isoform-level expression in single-cell populations. Our results reveal a common occurrence of isoform-level preference, commitment and heterogeneity in single-cell populations.

Bioinformatics

Simple multi-trait analysis identifies novel loci associated with growth and obesity measures

The ever-growing genome-wide association studies (GWAS) have revealed widespread pleiotropy. To exploit this, various methods which consider variant association with multiple traits jointly have been developed. However, most effort has been put on improving discovery power: how to replicate and interpret these discovered pleiotropic loci using multivariate methods has yet to be discussed fully. Using only multiple publicly available single-trait GWAS summary statistics, we develop a fast and flexible multi-trait framework that contains modules for (i) multi-trait genetic discovery, (ii) replication of locus pleiotropic profile, and (iii) multi-trait conditional analysis. The procedure is able to handle any level of sample overlap. As an empirical example, we discovered and replicated 23 novel pleiotropic loci for human anthropometry and evaluated their pleiotropic effects on other traits. By applying conditional multivariate analysis on the 23 loci, we discovered and replicated two additional multi-trait associated SNPs. Our results provide empirical evidence that multi-trait analysis allows detection of additional, replicable, highly pleiotropic genetic associations without genotyping additional individuals. The methods are implemented in a free and open source R package MultiABEL.\n\nAuthor summaryBy analyzing large-scale genomic data, geneticists have revealed widespread pleiotropy, i.e. single genetic variation can affect a wide range of complex traits. Methods have been developed to discover such genetic variants. However, we still lack insights into the relevant genetic architecture - What more can we learn from knowing the effects of these genetic variants?\n\nHere, we develop a fast and flexible statistical analysis procedure that includes discovery, replication, and interpretation of pleiotropic effects. The whole analysis pipeline only requires established genetic association study results. We also provide the mathematical theory behind the pleiotropic genetic effects testing.\n\nMost importantly, we show how a replication study can be essential to reveal new biology rather than solely increasing sample size in current genomic studies. For instance, we show that, using our proposed replication strategy, we can detect the difference in genetic effects between studies of different geographical origins.\n\nWe applied the method to the GIANT consortium anthropometric traits to discover new genetic associations, replicated in the UK Biobank, and provided important new insights into growth and obesity.\n\nOur pipeline is implemented in an open-source R package MultiABEL, sufficiently efficient that allows researchers to immediately apply on personal computers in minutes.

Genetics

Large-scale non-targeted metabolomic profiling in three human population-based studies

Metabolomic profiling is an emerging technique in life sciences. Human studies using these techniques have been performed in a small number of individuals or have been targeted at a restricted number of metabolites. In this article, we propose a data analysis workflow to perform non-targeted metabolomic profiling in large human population-based studies using ultra performance liquid chromatography-mass spectrometry (UPLC-MS). We describe challenges and propose solutions for quality control, statistical analysis and annotation of metabolic features. Using the data analysis workflow, we detected more than 8,000 metabolic features in serum samples from 2,489 fasting individuals. As an illustrative example, we performed a non-targeted metabolome-wide association analysis of high-sensitive C-reactive protein (hsCRP) and detected 407 metabolic features corresponding to 90 unique metabolites that could be replicated in an external population. Our results reveal unexpected biological associations, such as metabolites identified as monoacylphosphorylcholines (LysoPC) being negatively associated with hsCRP. R code and fragmentation spectra for all metabolites are made publically available. In conclusion, the results presented here illustrate the viability and potential of non-targeted metabolomic profiling in large population-based studies.

Bioinformatics