bioRxiv ScienceSearch

Biology subjects

Piepho, H.-P.

Publications and source records attributed to Piepho, H.-P..

2 recordsLinked to original sources

Effective principal components analysis of SNP data

PCA is frequently used to display and discover patterns in SNP data from humans, animals, plants, and microbes--especially to elucidate population structure. Given the popularity of PCA, one might expect that PCA is understood well and applied effectively. However, our literature survey of 125 representative articles that apply PCA to SNP data shows that three choices have usually been made poorly: SNP coding, PCA variant, and PCA graph. Accordingly, we offer several simple recommendations for effective PCA analysis of SNP data. The ultimate benefit from informed and optimal choices of SNP coding, PCA variant, and PCA graph is expected to be discovery of more biology, and thereby acceleration of medical, agricultural, and other vital applications.

genomics

Efficiency Of Genomic Prediction Of Non-Assessed Single Crosses

An important application of genomic selection in plant breeding is the prediction of untested single crosses (SCs). Most investigations on the prediction efficiency were based on tested SCs, using cross-validation. The main objective was to assess the prediction efficiency by correlating the predicted and true genotypic values of untested SCs (accuracy) and measuring the efficacy of identification of the best 300 untested SCs (coincidence), using simulated data. We assumed 10,000 SNPs, 400 QTLs, two groups of 70 selected DH lines, and 4,900 SCs. The heritabilities for the assessed SCs were 30, 60 and 100%. The scenarios included three sampling processes of DH lines, two sampling processes of SCs for testing, two SNP densities, DH lines from distinct and same populations, DH lines from populations with lower LD, two genetic models, three statistical models, and three statistical approaches. We derived a model for genomic prediction based on SNP average effects of substitution and dominance deviations. The prediction accuracy is not affected by the linkage phase. The prediction of untested SCs is very efficient. The accuracies and coincidences ranged from approximately 0.8 and 0.5, respectively, under low heritability, to 0.9 and 0.7, assuming high heritability. Additionally, we highlighted the relevance of the overall LD and evidenced that efficient prediction of untested SCs can be achieved for crops that show no heterotic pattern, for reduced training set size (10%), for SNP density of 1 cM, and for distinct sampling processes of DH lines, based on random choice of the SCs for testing.

genetics