bioRxiv Science⌕ Search

Biology subjects

Zhang, Y. D.

Publications and source records attributed to Zhang, Y. D..

2 recordsLinked to original sources

Variable Selection via Grace-AKO with Applications in Genomics

MotivationVariable selection is a common statistical approach to identifying genes associated with clinical outcomes of scientific interest. There are thousands of genes in genomic studies, while only a limited number of individual samples are available. Therefore, it is important to develop a method to identify genes associated with outcomes of interest that can control finite-sample false discovery rate (FDR) in high-dimensional data settings. ResultsThis article proposes a novel method named Grace-AKO for graph-constrained estimation (Grace), which incorporates aggregation of multiple knockoffs (AKO) with the network-constrained penalty. Grace-AKO can control FDR in finite-sample settings and improve model stability simultaneously. Simulation studies show that Grace-AKO has better performance in finite-sample FDR control than the original Grace model. We apply Grace-AKO to the prostate cancer data in The Cancer Genome Atlas (TCGA) program by incorporating prostate-specific antigen (PSA) pathways in the Kyoto Encyclopedia of Genes and Genomes (KEGG) as the prior information. Grace-AKO finally identifies 47 candidate genes associated with PSA level, and more than 75% of the detected genes can be validated. Availability and implementationWe developed an R package for Grace-AKO available at: https://github.com/mxxptian/GraceAKO Contactdoraz@hku.hk or zl2509@cumc.columbia.edu

bioinformatics↗

Multiethnic Polygenic Risk Prediction in Diverse Populations through Transfer Learning

Polygenic risk scores (PRS) leverage the genetic contribution of an individuals genotype to a complex trait by estimating disease risk. Traditional PRS prediction methods are predominantly for European population. The accuracy of PRS prediction in non-European populations is diminished due to much smaller sample size of genome-wide association studies (GWAS). In this article, we introduced a novel method to construct PRS for non-European populations, abbreviated as TL-Multi, by conducting transfer learning framework to learn useful knowledge from European population to correct the bias for non-European populations. We considered non-European GWAS data as the target data and European GWAS data as the informative auxiliary data. TL-Multi borrows useful information from the auxiliary data to improve the learning accuracy of the target data while preserving the efficiency and accuracy. To demonstrate the practical applicability of the proposed method, we applied TL-Multi to predict the risk of systemic lupus erythematosus (SLE) in Asian population and the risk of asthma in Indian population by borrowing information from European population. TL-Multi achieved better prediction accuracy than the competing methods including Lassosum and meta-analysis in both simulations and real applications.

genetics↗