bioRxiv Science⌕ Search

Biology subjects

Nguyen Huy, T.

Publications and source records attributed to Nguyen Huy, T..

3 recordsLinked to original sources

An efficient deep learning method for amino acid substitution model selection

Amino acid substitution models play an important role in studying the evolutionary relationships among species from protein sequences. The amino acid substitution model consists of a large number of parameters; therefore, it is estimated from hundreds or thousands of alignments. Both general models and clade-specific models have been estimated and widely used in phylogenetic analyses. The maximum likelihood method is normally used to select the best fit model for a specific protein alignment under the study. A number of studies have discussed theoretical concerns as well as computational burden of the maximum likelihood methods in model selection. Recently, machine learning methods have been proposed for selecting nucleotide models. In this paper, we propose methods to create summary statistics from protein alignments to efficiently train a network of so-called ModelDetector based on the convolutional neural network ResNet-18 for detecting amino acid models. Experiments on simulation data showed that the accuracy of ModelDetector was comparable with that of the maximum likelihood method ModelFinder. The ModelDetector network was trained from 64,800 alignments on a computer with 8 cores (without GPU) in about 12 hours. It is orders of magnitudes faster than the maximum likelihood method in inferring amino acid substitution models and able to analyze genome alignments with million sites in minutes.

bioinformatics↗

An Efficient Computational Method to Create Positive NIPT Samples with Autosomal Trisomy

Noninvasive prenatal test (NIPT) has been widely used for screening trisomy on chromosomes 13 (T13), 18 (T18), and 21 (T21). However, the false negative rate of NIPT algorithms has not been thoroughly evaluated due to the lack of positive samples. In this study, we present an efficient computational approach to create positive samples with autosomal trisomy from negative samples. We applied the approach to establish a low coverage dataset of 1440 positive samples with T13, T18, and T21 aberrations for both mosaic and non-mosaic conditions. We examined the performance of WisecondorX and its improvement, called VINIPT, on both negative and positive datasets. Experiments showed that WisecondorX and VINIPT with a z-score threshold of 3.3 were able to detect all non-mosaic samples with T13, T18, and T21 aberrations (i.e., the sensitivity of 100%). Using a lower z-score threshold of 2.58 when analyzing mosaic samples, both WisecondorX and VINIPT have the overall sensitivity of 99.7% on detecting T13, T18, and T21 aberrations from mosaic samples. WisecondorX has the specificity of 98.5% for the non-mosaic analysis, but a considerably lower specificity of 95% for the mosaic analysis. VINIPT has a much better specificity than WisecondorX, i.e., 99.9% for the non-mosaic analysis, and 98.2% for the mosaic analysis. The results suggest that VINIPT can play as a powerful NIPT tool for the low coverage data.

bioinformatics↗

Estimating amino acid substitution models from genome datasets: A simulation study on the performance of estimated models

Estimating amino acid substitution models is a crucial task in bioinformatics. The maximum likelihood (ML) approach has been proposed to estimate amino acid substitution models from large datasets. The quality of newly estimated models is normally assessed by comparing with the existing models in building ML trees. Two important questions remained are the correlation of the estimated models with the true models and the required size of the training datasets to estimate reliable models. In this paper, we performed a simulation study to answer these two questions based on the simulated data. We simulated genome datasets with different number of genes/alignments based on predefined models (called true models) and predefined trees (called true trees). The simulated datasets were used to estimate amino acid substitution model using the ML estimation method. Our experiment showed that models estimated by the ML methods from simulated datasets with more than 100 genes have high correlations with the true models. The estimated models performed well in building ML trees in comparison with the true models. The results suggest that amino acid substitution models estimated by the ML methods from large genome datasets might play as reliable tool for analyzing amino acid sequences.

bioinformatics↗