bioRxiv ScienceSearch

Biology subjects

Azencott, C.-A.

Publications and source records attributed to Azencott, C.-A..

2 recordsLinked to original sources

Novel Methods for Epistasis Detection in Genome-Wide Association Studies

More and more genome-wide association studies are being designed to uncover the full genetic basis of common diseases. Nonetheless, the resulting loci are often insufficient to fully recover the observed heritability. Epistasis, or gene-gene interaction, is one of many hypotheses put forward to explain this missing heritability. In the present work, we propose epiGWAS, a new approach for epistasis detection that identifies interactions between a target SNP and the rest of the genome. This contrasts with the classical strategy of epistasis detection through exhaustive pairwise SNP testing. We draw inspiration from causal inference in randomized clinical trials, which allows us to take into account linkage disequilibrium. EpiGWAS encompasses several methods, which we compare to state-of-the-art techniques for epistasis detection on simulated and real data. The promising results demonstrate empirically the benefits of EpiGWAS to identify pairwise interactions. Author summaryGenome-wide association studies are now a major tool for the discovery of biomarkers for complex diseases. However, the complexity of genetic architecture, in particular linkage disequilibrium, complicates that mission. Moreover, intergenic interactions, or epistasis, are often not correctly captured by the classical statistical methodologies. In our work, we propose a new framework to model linkage disequilibrium, which is based on propensity scores. Our goal is to detect epistatic interactions between a predetermined target locus and the rest of the genotype. The target may be identified from the literature, experiments, or top hits in previous genome-wide association studies. Recovering interactions with validated causal loci helps improve both interpretability and statistical power. Multi-targeting drug discovery can also benefit from our work through the combination of existing drugs with new ones for greater drug response.

bioinformatics

Efficient Multi-task chemogenomics for drug specificity prediction

Adverse drug reactions, also called side effects, range from mild to fatal clinical events and significantly affect the quality of care. Among other causes, side effects occur when drugs bind to proteins other than their intended target. As experimentally testing drug specificity against the entire proteome is out of reach, we investigate the application of chemogenomics approaches. We formulate the study of drug specificity as a problem of predicting interactions between drugs and proteins at the proteome scale. We build several benchmark datasets, and propose NN-MT, a multi-task Support Vector Machine (SVM) algorithm that is trained on a limited number of data points, in order to solve the computational issues or proteome-wide SVM for chemogenomics. We compare NN-MT to different state-of-the-art methods, and show that its prediction performances are similar or better, at an efficient calculation cost. Compared to its competitors, the proposed method is particularly efficient to predict (protein, ligand) interactions in the difficult double-orphan case, i.e. when no interactions are previously known for the protein nor for the ligand. The NN-MT algorithm appears to be a good default method providing state-of-the-art or better performances, in a wide range of prediction scenarii that are considered in the present study: proteome-wide prediction, protein family prediction, test (protein, ligand) pairs dissimilar to pairs in the train set, and orphan cases.

bioinformatics