bioRxiv ScienceSearch

Biology subjects

Prive, F.

Publications and source records attributed to Prive, F..

4 recordsLinked to original sources

Efficient implementation of penalized regression for genetic risk prediction

Polygenic Risk Scores (PRS) consist in combining the information across many single-nucleotide polymorphisms (SNPs) in a score reflecting the genetic risk of developing a disease. PRS might have a major impact on public health, possibly allowing for screening campaigns to identify high-genetic risk individuals for a given disease. The \"Clumping+Thresholding\" (C+T) approach is the most common method to derive PRS. C+T uses only univariate genome-wide association studies (GWAS) summary statistics, which makes it fast and easy to use. However, previous work showed that jointly estimating SNP effects for computing PRS has the potential to significantly improve the predictive performance of PRS as compared to C+T.\n\nIn this paper, we present an efficient method to jointly estimate SNP effects, allowing for practical application of penalized logistic regression (PLR) on modern datasets including hundreds of thousands of individuals. Moreover, our implementation of PLR directly includes automatic choices for hyper-parameters. The choice of hyper-parameters for a predictive model is very important since it can dramatically impact its predictive performance. As an example, AUC values range from less than 60% to 90% in a model with 30 causal SNPs, depending on the p-value threshold in C+T.\n\nWe compare the performance of PLR, C+T and a derivation of random forests using both real and simulated data. PLR consistently achieves higher predictive performance than the two other methods while being as fast as C+T. We find that improvement in predictive performance is more pronounced when there are few effects located in nearby genomic regions with correlated SNPs; for instance, AUC values increase from 83% with the best prediction of C+T to 92.5% with PLR. We confirm these results in a data analysis of a case-control study for celiac disease where PLR and the standard C+T method achieve AUC of 89% and of 82.5%.\n\nIn conclusion, our study demonstrates that penalized logistic regression can achieve more discriminative polygenic risk scores, while being applicable to large-scale individual-level data thanks to the implementation we provide in the R package bigstatsr.

genetics

VarExp: Estimating variance explained by Genome-Wide GxE summary statistics

Many genomic analyses, such as genome-wide association studies (GWAS) or genome-wide screening for Gene-Environment (GxE) interactions have been performed to elucidate the underlying mechanisms of human traits and diseases. When the analyzed outcome is quantitative, the overall contribution of identified genetic variants to the outcome is often expressed as the percentage of phenotypic variance explained. In practice, this is commonly estimated using individual genotype data. However, using individual-level data faces practical and ethical challenges when the GWAS results are derived in large consortia through meta-analysis of results from multiple cohorts. In this work, we present a R package, \"VarExp\", that allows for the estimation of the percentage of phenotypic variance explained by variants of interest using summary statistics only. Our package allows for a range of models to be evaluated, including marginal genetic effects, GxE interaction effects, and main genetic and interaction effects jointly. Its implementation integrates all recent methodological developments on the topic and does not need external data to be uploaded by users.\n\nThe R source code, tutorial and associated example are available at https://gitlab.pasteur.fr/statistical-genetics/VarExp.git.

bioinformatics

Efficient management and analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr

MotivationGenome-wide datasets produced for association studies have dramatically increased in size over the past few years, with modern datasets commonly including millions of variants measured in dozens of thousands of individuals. This increase in data size is a major challenge severely slowing down genomic analyses. Specialized software for every part of the analysis pipeline have been developed to handle large genomic data. However, combining all these software into a single data analysis pipeline might be technically difficult.\n\nResultsHere we present two R packages, bigstatsr and bigsnpr, allowing for management and analysis of large scale genomic data to be performed within a single comprehensive framework. To address large data size, the packages use memory-mapping for accessing data matrices stored on disk instead of in RAM. To perform data pre-processing and data analysis, the packages integrate most of the tools that are commonly used, either through transparent system calls to existing software, or through updated or improved implementation of existing methods. In particular, the packages implement a fast derivation of Principal Component Analysis, functions to remove SNPs in Linkage Disequilibrium, and algorithms to learn Polygenic Risk Scores on millions of SNPs. We illustrate applications of the two R packages by analysing a case-control genomic dataset for the celiac disease, performing an association study and computing Polygenic Risk Scores. Finally, we demonstrate the scalability of the R packages by analyzing a simulated genome-wide dataset including 500,000 individuals and 1 million markers on a single desktop computer.\n\nAvailabilityhttps://privefl.github.io/bigstatsr/ & https://privefl.github.io/bigsnpr/\n\nContactflorian.prive@univ-grenoble-alpes.fr & michael.blum@univ-grenoble-alpes.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

The rumen microbiome: an underexplored resource for novel antimicrobial discovery

Antimicrobial peptides (AMPs) are promising drug candidates to target multi-drug resistant bacteria. The rumen microbiome presents an underexplored resource for the discovery of novel microbial enzymes and metabolites, including AMPs. Using functional screening and computational approaches, we identified 181 potentially novel AMPs from a rumen bacterial metagenome. Here, we show that three of the selected AMPs (Lynronne-1, 2 and 3) were effective against numerous bacterial pathogens, including methicillin resistant Staphylococcus aureus (MRSA). No decrease in MRSA susceptibility was observed after 25 days of sub-lethal exposure to these AMPs. The AMPs bound preferentially to bacterial membrane lipids and induced membrane permeability leading to cytoplasmic leakage. Topical administration of Lynronne-1 (10% w/v) to a mouse model of MRSA wound infection elicited a significant reduction in bacterial counts, which was comparable to treatment with 2% mupirocin ointment. Our findings indicate that the rumen microbiome may provide viable alternative antimicrobials for future therapeutic application.

microbiology