bioRxiv Science⌕ Search

Biology subjects

Slomian, D.

Publications and source records attributed to Slomian, D..

2 recordsLinked to original sources

Modeling missing parents in single-step test-day SNP-BLUP evaluation of dairy cattle

In many countries, single-step genomic models have replaced multiple step models for routine evaluation. These models use all available information on animals phenotypes, genotypes, and pedigrees, yet missing parental information in pedigrees remains a challenge that affects genomic breeding value (GEBV) predictions. Therefore, the choice of method for handling missing parents can affect the prediction of breeding values. Here, we compared three approaches to model missing parental information for three levels of missing pedigree data: P_Real - pedigree from routine evaluation, P_2010 - at least 20 percent of dams and 10 percent of sires born before 2019 were set to missing, and P_4020 - at least 40 percent of dams and 20 percent of sires born before 2019 were set to missing. Missing parents information was expressed through missing codes in the raw pedigree (RP) by defining genetic groups (GG) that represent missing parents grouped based on year of birth, sex, and country of origin, or by defining metafounders (MF), which represent missing parents grouped by average genetic relationships estimated from the genomic information of their descendants. The genomic breeding values for fat yield were estimated using the single-step test-day SNP-BLUP model implemented with MiXBLUP software. For the considered scenarios, the results were presented separately for sires and dams, as well as for genotyped and ungenotyped individuals. We observed differences in the prediction quality between genotyped and ungenotyped animals. While GEBV predictions for the former were generally stable across scenarios, the predictions for the ungenotyped individuals varied. In particular, the removal of parental information led to less stable results when missing parental information was expressed by MF, where insufficient pedigree completeness resulted in an overestimation of the genetic trend. In conclusion, for informative pedigrees with a small percentage of missing parents, the incorporation of GG and MF results in very similar GEBV predictions, however GG appear to be a more robust approach for ungenotyped individuals in highly incomplete pedigrees.

animal behavior and cognition↗

Approaches to dimensionality reduction for ultra-high dimensional models

The rapid advancement of high-throughput sequencing technologies has revolutionised genomic research by providing access to large amounts of genomic data. However, the most important disadvantage of using Whole Genome Sequencing (WGS) data is its statistical nature, the so-called p>>n problem. This study aimed to compare three approaches of feature selection allowing for circumventing the p>>n problem, among which one is a novel modification of Supervised Rank Aggregation (SRA). The use of the three methods was demonstrated in the classification of 1,825 individuals representing the 1000 Bull Genomes Project to 5 breeds, based on 11,915,233 SNP genotypes from WGS. In the first step, we applied three feature (i.e. SNP) selection methods: the mechanistic approach (SNP tagging) and two approaches considering biological and statistical contexts by fitting a multiclass logistic regression model followed by either 1-dimensional clustering (1D-SRA) or multi-dimensional feature clustering (MD-SRA) that was originally proposed in this study. Next, we perform the classification based on a Deep Learning architecture composed of Convolutional Neural Networks. The classification quality of the test data set was expressed by macro F1-Score. The SNPs selected by SNP tagging yielded the least satisfactory results (86.87%). Still, this approach offered rapid computing times by focussing only on pairwise LD between SNPs and disregarding the effects of SNP on classification. 1D-SRA was less suitable for ultra-high-dimensional applications due to computational, memory and storage limitations, however, the SNP set selected by this approach provided the best classification quality (96.81%). MD-SRA provided a very good balance between classification quality (95.12%) and computational efficiency (17x lower analysis time and 14x lower data storage), outperforming other methods. Moreover, unlike SNP tagging, both SRA-based approaches are universal and not limited to feature selection for genomic data. Our work addresses the urgent need for computational techniques that are both effective and efficient in the analysis and interpretation of large-scale genomic datasets. We offer a model suitable for the classification of ultra-high-dimensional data that implements fusing feature selection and deep learning techniques.

bioinformatics↗