bioRxiv Science⌕ Search

Biology subjects

Borin, V. A.

Publications and source records attributed to Borin, V. A..

3 recordsLinked to original sources

Taxonomic Resolution of 16S rRNA, FastANI, Mash, and FastAAI across 30,495 Prokaryotic Type-Strain Genomes

Prokaryotic taxonomy now relies on both marker-gene and genome-wide sequence comparisons, but these methods differ in taxonomic range, scalability, and sensitivity to genome quality. Here, we benchmarked four commonly used approaches, including 16S rRNA identity, FastANI, Mash distance, and FastAAI/Jaccard similarity across a dataset of 30,495 prokaryotic type-strain genomes. Type-strain genomes provide nomenclatural anchors for validly named species, making them a useful framework for evaluating how sequence-based methods correspond to current taxonomic assignments. We evaluated method behavior across taxonomic ranks from species to domain and separated initial method failures from threshold-based failures. When clean full-length 16S rRNA sequences were available, same-species comparisons passed the empirical threshold in >97% of cases. However, a usable full-length 16S rRNA sequence was unavailable for 4,551 of the 16,402 same-species comparisons (28%), limiting marker-gene-based analysis. In addition, 16S rRNA identity ranges overlapped across higher taxonomic ranks, limiting the use of universal rank-specific cutoffs. FastANI provided strong species-level resolution, with same-species comparisons passing the empirical threshold in approximately 88% of cases but was less informative at deeper ranks. Mash enabled rapid genome-scale screening, although its distance values require careful interpretation beyond close relatives. FastAAI provided a genome-wide amino-acid signal, with approximately 92% of same-species comparisons passing the empirical threshold and was especially useful for comparisons beyond the species boundary. Overall, no single method performed optimally across all taxonomic levels. These results support a rank-aware benchmarking framework in which 16S rRNA, FastANI, Mash, and FastAAI are interpreted as complementary tools, with attention to genome quality, missing data, and method-specific failure modes.

genomics↗

APE1 active site residue Asn174 stabilizes the AP-site and is essential for catalysis

Apurinic/Apyrimidinic (AP)-sites are common and highly mutagenic DNA lesions that can arise spontaneously or as intermediates during Base Excision Repair (BER). The enzyme apurinic/apyrimidinic endonuclease 1 (APE1) initiates repair of AP-sites by cleaving the DNA backbone at the AP-site via its endonuclease activity. Here, we investigated the functional role of the APE1 active site residue N174 that contacts the AP-site during catalysis. We analyzed the effects of three rationally designed APE1 mutations that alter the hydrogen bonding potential, size, and charge of N174: N174A, N174D, and N174Q. We found impaired catalysis of the APE1N174A and APE1N174D mutants due to disruption of hydrogen bonding and electrostatic interactions between residue 174 and the AP-site. In comparison, the APE1N174Q mutant was less impaired due to retaining similar hydrogen bonding and electrostatic characteristics as N174 in wild-type APE1. Structures and computational simulations further revealed that the AP-site was destabilized within the active sites of the APE1N174A and APE1N174D mutants due to loss of hydrogen bonding between residue 174 and the AP-site. Cumulatively, we show that N174 stabilizes the AP-site within the APE1 active site through hydrogen bonding and electrostatic interactions to enable effective catalysis. These findings highlight the importance of N174 in APE1s function and provide new insights into the molecular mechanism by which APE1 processes AP-sites during DNA repair.

biochemistry↗

Pandemic preparedness through genomic surveillance: Overview of mutations in SARS-CoV-2 over the course of COVID-19 outbreak

Genomic surveillance is a vital strategy for preparedness against the spread of infectious diseases and to aid in development of new treatments. In an unprecedented effort, millions of samples from COVID-19 patients have been sequenced worldwide for SARS-CoV-2. Using more than 8 million sequences that are currently available in GenBanks SARS-CoV-2 database, we report a comprehensive overview of mutations in all 26 proteins and open reading frames (ORFs) from the virus. The results indicate that the spike protein, NSP6, nucleocapsid protein, envelope protein and ORF7b have shown the highest mutational propensities so far (in that order). In particular, the spike protein has shown rapid acceleration in mutations in the post-vaccination period. Monitoring the rate of non-synonymous mutations (Ka) provides a fairly reliable signal for genomic surveillance, successfully predicting surges in 2022. Further, the external proteins (spike, membrane, envelope, and nucleocapsid proteins) show a significant number of mutations compared to the NSPs. Interestingly, these four proteins showed significant changes in Ka typically 2 to 4 weeks before the increase in number of human infections ("surges"). Therefore, our analysis provides real time surveillance of mutations of SARS-CoV-2, accessible through the project website http://pandemics.okstate.edu/covid19/. Based on ongoing mutation trends of the virus, predictions of what proteins are likely to mutate next are also made possible by our approach. The proposed framework is general and is thus applicable to other pathogens. The approach is fully automated and provides the needed genomic surveillance to address a fast-moving pandemic such as COVID-19.

genomics↗