bioRxiv Science⌕ Search

Biology subjects

Dekkers, J. C. M.

Publications and source records attributed to Dekkers, J. C. M..

3 recordsLinked to original sources

Using encrypted genotypes and phenotypes for collaborative genomic analyses to maintain data confidentiality

To adhere to and capitalize on the benefits of the FAIR (Findable, Accessible, Interoperable and Reusable) principles in agricultural genome-to-phenome studies, it is crucial to address privacy and intellectual property issues that prevent sharing and reuse of data in research and industry. Direct sharing of genotype and phenotype data is often prohibited due to intellectual property and privacy concerns. Thus there is a pressing need for encryption methods that obscure confidential aspects of the data, without affecting the outcomes of certain statistical analyses. A homomorphic encryption method for genotypes and phenotypes (HEGP) has been proposed for single-marker regression in genome-wide association studies using linear mixed models with Gaussian errors. This methodology permits frequentist likelihood-based parameter estimation and inference. In this paper, we extend HEGP to broader applications in genome-to-phenome analyses. We show that HEGP is suited to commonly used linear mixed models for genetic analyses of quantitative traits including GBLUP and RR-BLUP, as well as Bayesian variable selection methods (e.g., those in Bayesian Alphabet), for genetic parameter estimation, genomic prediction, and genome-wide association studies. By advancing the capabilities of HEGP, we offer researchers and industry professionals a secure and efficient approach for collaborative genomic analyses while preserving data confidentiality.

genetics↗

A compendium of genetic regulatory effects across pig tissues

The Farm animal Genotype-Tissue Expression (FarmGTEx, https://www.farmgtex.org/) project has been established to develop a comprehensive public resource of genetic regulatory variants in domestic animal species, which is essential for linking genetic polymorphisms to variation in phenotypes, helping fundamental biology discovery and exploitation in animal breeding and human biomedicine. Here we present results from the pilot phase of PigGTEx (http://piggtex.farmgtex.org/), where we processed 9,530 RNA-sequencing and 1,602 whole-genome sequencing samples from pigs. We build a pig genotype imputation panel, characterize the transcriptional landscape across over 100 tissues, and associate millions of genetic variants with five types of transcriptomic phenotypes in 34 tissues. We study interactions between genotype and breed/cell type, evaluate tissue specificity of regulatory effects, and elucidate the molecular mechanisms of their action using multi-omics data. Leveraging this resource, we decipher regulatory mechanisms underlying about 80% of the genetic associations for 207 pig complex phenotypes, and demonstrate the similarity of pigs to humans in gene expression and the genetic regulation behind complex phenotypes, corroborating the importance of pigs as a human biomedical model.

genomics↗

Validation of the linear regression method to evaluate population accuracy and bias of predictions for non-linear models

BackgroundThe linear regression method (LR) was proposed to estimate population bias and accuracy of predictions, while addressing the limitations of commonly used cross-validation methods. The validity and behavior of the LR method have been provided and studied for linear model predictions but not for non-linear models. The objectives of this study were to 1) provide a mathematical proof for the validity of the LR method when predictions are based on conditional mean, 2) explore the behavior of the LR method in estimating bias and accuracy of predictions when the model fitted is different from the true model, and 3) provide guidelines on how to appropriately partition the data into training and validation such that the LR method can identify presence of bias and accuracy in predictions. ResultsWe present a mathematical proof for the validity of the LR method to estimate bias and accuracy of predictions based on the conditional mean, including for non-linear models. Using simulated data, we show that the LR method can accurately detect bias and estimate accuracy of predictions when an incorrect model is fitted when the data is partitioned such that the values of relevant predictor variables differ in the training and validation sets. But the LR method fails when the data are not partitioned in that manner. ConclusionsThe LR method was proven to be a valid method to evaluate the population bias and accuracy of predictions based on the conditional mean, regardless of whether it is a linear or non-linear function of the data. The ability of the LR method to detect bias and estimate accuracy of predictions when the model fitted is incorrect depends on how the data are partitioned. To appropriately test the predictive ability of a model using the LR method, the values of the relevant predictor variables need to be different between the training and validation sets.

genetics↗