bioRxiv Science⌕ Search

Biology subjects

Kangzhu, Y.

Publications and source records attributed to Kangzhu, Y..

2 recordsLinked to original sources

KBeagle: An Adaptive Strategy and Tool for Improvement of Imputation Accuracy and Computing Efficiency

With the development of molecular biology and genetics, deep sequencing technology has become the main way to discover genetic variation and reveal the molecular structure of genome. Due to the complexity of the whole genome segment structure, a large number of missing genotypes have appeared after sequencing, and these missing genotypes can be imputed by genotype imputation method. With the in-depth study of genotype imputation methods, computational intensive and computationally efficient imputation software come into being. Beagle software, as an efficient imputation software, is widely used because of its advantages of low memory consumption, fast running speed and relatively high imputation accuracy. K-Means clustering can divide individuals with similar population structure into a class, so that individuals in the same class can share longer haplotype fragments. Therefore, combining K-Means clustering algorithm with Beagle software can improve the interpolation accuracy. The Beagle and KBeagle method was used to compare the imputation efficiency. The KBeagle method presents a higher imputation matching rate and a shorter computing time. In the genome selection and heritability estimated section, the genotype dataset after imputed, unimputed, and with real genotype show similar prediction accuracy. However the estimated heritability using genotype dataset after imputed is closer to the estimation by the dataset with real genotype. We generated a compounds and efficient imputation method, which presents valuable resource for improvement of imputation accuracy and computing time. We envisage the application of KBeagle will be focus on the livestock sequencing study under strong genetic structure.

bioinformatics↗

pCMLM: Genome Wide Association Study of Body Size Traits in Multiple Regions of Yak Based on the Provided Compressed Mixed Linear Model

ObjectiveYak is a unique large animal species living in the Qinghai-Tibet Plateau and the surrounding Hengduan Mountains, and has evolved several regional variety resources due to the special geographical and ecological environment in which it lives. Therefore, it is of great importance to investigate the genetic composition of body size traits among breeds in multiple regions for yak breeding and production. MethodA genome-wide association analysis was performed on 94 yak individuals (a total of 31 variety resources) for five body size traits (body height, body weight, body length, chest circumference, and circumference of cannon bone). The individuals were clustered following known population habitat. The kinship of grouping individuals was used in the CMLM. This provided compressed mixed linear model was named pCMLM method. ResultTotal of 3,584,464 high-quality SNP markers were obtained on 30 chromosomes. Principal component analysis using the whole SNPs do not accurately classify all populations into multiple subpopulations, a result that is not the same as the population habitat. Six SNP loci were identified in the pCMLM-based GWAS with statistically significant correlation with body height, and four candidate genes (FXYD6, SOHLH2, ADGRB2, and OSBPL6), which in the vicinity of the variant loci, were screened and annotated. Two of these genes, ADGRB2 and OSBPL6, are involved in biological regulatory processes such as body height regulation, adipocyte proliferation and differentiation. ConclusionBased on the previous population information, the pCMLM can provide more sufficient associated results when the conventional CMLM can not catch optimum clustering groups. The fundamental information for quantitative trait gene localization or candidate gene cloning in the mechanism of yak body size trait formation.

bioinformatics↗