bioRxiv ScienceSearch

Biology subjects

Chu, Y.

Publications and source records attributed to Chu, Y..

2 recordsLinked to original sources

Machine learning as an effective method for identifying true SNPs in polyploid plants

Single Nucleotide Polymorphisms (SNPs) have many advantages as molecular markers since they are ubiquitous and co-dominant. However, the discovery of true SNPs especially in polyploid species is difficult. Peanut is an allopolyploid, which has a very low rate of true SNP calling. A large set of true and false SNPs identified from the Arachis 58k Affymetrix array was leveraged to train machine learning models to select true SNPs straight from sequence data. These models achieved accuracy rates of above 80% using real peanut RNA-seq and whole genome shotgun (WGS) re-sequencing data, which is higher than previously reported for polyploids. A 48K SNP array, Axiom Arachis2, was designed using the approach which revealed 75% accuracy of calling SNPs from different tetraploid peanut genotypes. Using the method to simulate SNP variation in peanut, cotton, wheat, and strawberry, we show that models built with our parameter sets achieve above 98% accuracy in selecting true SNPs. Additionally, models built with simulated genotypes were able to select true SNPs at above 80% accuracy using real peanut data, demonstrating that our model can be used even if real data are not available to train the models. This work demonstrates an effective approach for calling highly reliable SNPs from polyploids using machine learning. A novel tool was developed for predicting true SNPs from sequence data, designated as SNP-ML (SNP-Machine Learning, pronounced \"snip mill\"), using the described models. SNP-ML additionally provides functionality to train new models not included in this study for customized use, designated SNP-MLer (SNP-Machine Learner, pronounced \"snip miller\"). SNP-ML is freely available for public use.

bioinformatics

Insight Into The Mechanism Of Protein Thermostability Based On The Residue Interaction Degrees

Understanding the basis of protein thermostability raises a general question: which residue with specific interaction degrees is more important to the protein thermostability? A strictly selected dataset of 131 pairs of thermophilic (TPs) and mesophilic proteins (MPs) was constructed. There were 6.4% and 8.4% of the total residues in sequences did not interact with others in TPs and MPs. The amino acid contents in sequences are closest to those with the interaction degrees of 3 according to the Chi-squared distances. Only Glu, Gln and the amide residues showed significant differences in sequences, which was the same as identified at low residue interaction degrees. However, we observed significant Phe, Lys, Leu, Gln and the charged, aliphatic, aromatic, positive charged and small residues at high interaction degree. Among them, Phe was rarely reported previously although aromatic residues were well-known contributor to protein thermostability. Finally, we took aspartate transcarbamylases as an example to explain how a residue with various interaction degrees contributed differently to their thermostability. Our results clearly demonstrated the differences of amino acids in sequence between TPs and MPs could only represent those involved in low interaction degrees. Much more residues with significant differences existed at high interaction degrees even if they had few significant amino acids in sequences. The interaction degree-based method should be an alternative tool in extracting valuable eigenvalues for predicting proteins attributes in bioinformatics. It could also provide a new perspective for studying the thermostability of proteins and engineering novel thermostable proteins.\n\nList of abbreviations

bioinformatics