bioRxiv Science⌕ Search

Biology subjects

Tsuruta, S.

Publications and source records attributed to Tsuruta, S..

2 recordsLinked to original sources

Marker effect p-values for single-step GWAS with the algorithm for proven and young in large genotyped populations

BackgroundAlthough single-step GBLUP (ssGBLUP) is a breeding value method, single-nucleotide polymorphism (SNP) effects can be backsolved from ssGBLUP genomic estimated breeding values (GEBV), and p-values can be obtained as a measure of estimation certainty. This enables single-step genome-wide association studies (ssGWAS). However, obtaining p-values for ssGWAS relies on the inversion of dense matrices, which poses computational limitations in large genotyped populations. In this study, we present an algorithm to approximate p-values for SNP in ssGWAS with many genotyped animals. The approximation relies on the algorithm for proven and young (APY) and submatrices for core animals. To test that, we first compared SNP p-values obtained with an exact inversion using the genomic relationship matrix (G-1) for 50K genotyped animals to those estimated with an exact inversion using [Formula] and those obtained with the proposed approximation based on [Formula]. Then, we compared these results with those obtained with the proposed approximation using 450K genotyped animals. ResultsThe same genomic regions in chromosomes 7 and 20 were identified with p-values obtained with G-1, [Formula], and the approximation based on [Formula] when using 50k genotyped animals and 1.5M in the pedigree. In terms of computational requirements, obtaining p-values with the approximation based on [Formula] represented a reduction of 38 times in wall-clock time and ten times in memory requirement compared to using the exact inversion with [Formula]. When the approximation was applied to a population of 450K genotyped animals and 1.8 in the pedigree, apart from the two genomic regions in chromosomes 7 and 20 previously identified with the smaller genotyped population, two new significant regions on chromosomes 6 and 14 were uncovered, indicating an increase in GWAS detection power when including more genotypes in the analyses. The process of obtaining p-value with the approximation and 450K genotyped individuals took 24.5 wall-clock hours and 87.66 GB of memory, which is expected to increase linearly with the addition of noncore genotyped individuals. ConclusionsWith an algorithm that approximates the prediction error variance of SNP effects based on APY, ssGWAS with p-values for SNP is possible in large genotyped populations. The computational cost of obtaining p-values in ssGWAS is no longer a limitation in extensive populations with many genotyped animals.

genetics↗

Dimensionality of genomic information and its impact on GWA and variant selection: a simulation study

BackgroundIdentifying true-positive variants in genome-wide associations (GWA) depends on several factors, including the number of genotyped individuals. The limited dimensionality of the genomic information may give insights into the optimal number of individuals to use in GWA. This study investigated different discovery set sizes in GWA based on the number of largest eigenvalues explaining a certain proportion of variance in the genomic relationship matrix (G). An additional investigation included the change in accuracy by adding variants, selected based on different set sizes, to the regular SNP chips used for genomic prediction. MethodsSequence data were simulated containing 500k SNP with 200 or 2000 quantitative trait nucleotides (QTN). A regular 50k panel included one every ten simulated SNP. Effective population size (Ne) was 20 and 200. The GWA was performed with the number of genotyped animals equivalent to the number of largest eigenvalues of G (EIG) explaining 50, 60, 70, 80, 90, 95, 98, and 99% of the variance. In addition, the largest discovery set consisted of 30k genotyped animals. Limited or extensive phenotypic information was mimicked by changing the trait heritability. Significant and high effect size SNP were added to the 50k panel and used for single-step GBLUP with and without weights. ResultsUsing the number of genotyped animals corresponding to at least EIG98 enabled the identification of QTN with the largest effect sizes when Ne was large. Smaller populations required more than EIG98. Furthermore, using genotyped animals with higher reliability (i.e., higher trait heritability) helped better identify the most informative QTN. The greatest prediction accuracy was obtained when the significant or the high effect SNP representing twice the number of simulated QTN were added to the 50k panel. Weighting SNP differently did not increase prediction accuracy, mainly because of the size of the genotyped population. ConclusionsAccurately identifying causative variants from sequence data depends on the effective population size and, therefore, the dimensionality of genomic information. This dimensionality can help identify the suitable sample size for GWA and could be considered for variant selection. Even when variants are accurately identified, their inclusion in prediction models has limited implications.

genomics↗