bioRxiv Science⌕ Search

Biology subjects

Aw, A. J.

Publications and source records attributed to Aw, A. J..

4 recordsLinked to original sources

Robust and Adaptive Non-Parametric Tests for Detecting General Distributional Shifts in Gene Expression

Differential expression analysis is crucial in genomics, yet existing methods primarily focus on detecting mean shifts. Variance shifts in gene expression are well-documented in studies of cellular signaling pathways, and more recently they have characterized aging, thus motivating the need for flexible detection approaches that include tests of expression variance changes. In this work, we present QRscore (Quantile Rank Score), a general method for detecting distributional shifts in gene expression by extending the Mann-Whitney test into a flexible family of rank-based tests. Here, we focus on implementing QRscore to detect shifts in mean and variance in gene expression, using weights designed from negative binomial (NB) and zero-inflated negative binomial (ZINB) models to combine the strengths of parametric and non-parametric approaches. We show through simulations that QRscore not only achieves high statistical power while controlling the false discovery rate (FDR), but also outperforms existing methods in detecting variance shifts and mean shifts. Applying QRscore to bulk RNA-seq data from the Genotype-Tissue Expression (GTEx) project, we identified numerous differentially dispersed genes and differentially expressed genes across 33 tissues. Notably, many genes have significant variance shifts but non-significant mean shifts. QRscore augments the genome bioinformatics toolkit by offering a powerful and flexible approach for differential expression analysis. QRscore is available in R, at https://github.com/songlab-cal/QRscore.

bioinformatics↗

Highly parameterized polygenic scores tend to overfit to population stratification via random effects

Polygenic scores (PGSs), increasingly used in clinical settings, frequently include many genetic variants, with performance typically peaking at thousands of variants. Such highly parameterized PGSs often include variants that do not pass a genome-wide significance threshold. We propose a mathematical perspective that renders the effects of many of these nonsignificant variants random rather than causal, with the randomness capturing population structure. We devise methods to assess variant effect randomness and population stratification bias. Applying these methods to 141 traits from the UK Biobank, we find that, for many PGSs, the effects of non-significant variants are considerably random, with the extent of randomness associated with the degree of overfitting to population structure of the discovery cohort. Our findings explain why highly parameterized PGSs simultaneously have superior cohort-specific performance and limited generalizability, suggesting the critical need for variant randomness tests in PGS evaluation. Supporting code and a dashboard are available at https://github.com/songlab-cal/StratPGS.

genetics↗

GPN-MSA: an alignment-based DNA language model for genome-wide variant effect prediction

Whereas protein language models have demonstrated remarkable efficacy in predicting the effects of missense variants, DNA counterparts have not yet achieved a similar competitive edge for genome-wide variant effect predictions, especially in complex genomes such as that of humans. To address this challenge, we here introduce GPN-MSA, a novel framework for DNA language models that leverages whole-genome sequence alignments across multiple species and takes only a few hours to train. Across several benchmarks on clinical databases (ClinVar, COSMIC, OMIM), experimental functional assays (DMS, DepMap), and population genomic data (gnomAD), our model for the human genome achieves outstanding performance on deleteriousness prediction for both coding and non-coding variants.

bioinformatics↗

Comprehensive mapping of SARS-CoV-2 interactions in vivo reveals functional virus-host interactions

SARS-CoV-2 has emerged as a major threat to global public health, resulting in global societal and economic disruptions. Here, we investigate the intramolecular and intermolecular RNA interactions of wildtype (WT) and a mutant ({Delta}382) SARS-CoV-2 virus in cells using high throughput structure probing on Illumina and Nanopore platforms. We identified twelve potentially functional structural elements within the SARS-CoV-2 genome, observed that identical sequences can fold into divergent structures on different subgenomic RNAs, and that WT and {Delta}382 virus genomes can fold differently. Proximity ligation sequencing experiments identified hundreds of intramolecular and intermolecular pair-wise interactions within the virus genome and between virus and host RNAs. SARS-CoV-2 binds strongly to mitochondrial and small nucleolar RNAs and is extensively 2-O-methylated. 2-O-methylation sites in the virus genome are enriched in the untranslated regions and are associated with increased pair-wise interactions. SARS-CoV-2 infection results in a global decrease of 2-O-methylation sites on host mRNAs, suggesting that binding to snoRNAs could be a pro-viral mechanism to sequester methylation machinery from host RNAs towards the virus genome. Collectively, these studies deepen our understanding of the molecular basis of SARS-CoV-2 pathogenicity, cellular factors important during infection and provide a platform for targeted therapy.

genomics↗