bioRxiv Science⌕ Search

Biology subjects

San, J. E.

Publications and source records attributed to San, J. E..

2 recordsLinked to original sources

SeqPanther: Sequence manipulation and mutation statistics toolset

Pathogen genomes harbor critical information necessary to support genomic investigations that inform public health interventions such as treatment, control, and eradication. To extract this information, their sequences are analysed to identify structural variations such as single nucleotide polymorphisms (SNPs) and insertions and deletions (indels) that may be associated with phenotypes of interest. Typically, this involves generating a consensus sequence from raw reads, aligning it to a reference and identifying positions where variations occur. Several pipelines exist to map raw reads and assemble whole genomes for downstream analysis. However, there is no easy to use, freely available bioinformatics quality control (QC) tool to explore mappings for both positional codons and nucleotide distributions in mapped short reads of microbial genomes. To address this problem, we have developed a fast and accurate tool to summarise read counts associated with codons, nucleotides, and indels in mapped next-generation sequencing (NGS) short reads. The tool, developed in Python, also provides a visualization of the genome sequencing depth and coverage. Furthermore, the tool can be run in single or batch mode, where several genomes need to be analysed. Our tool produces a text-based report that enables quick review or can be imported into any analytical tool for upstream analysis. Additionally, the tool also provides functionality to modify the consensus sequences by adding, masking, or restoring to wild type mutations specified by the user. AvailabilitySeqPanther is available at https://github.com/codemeleon/seqPanther, along with the necessary documentation for installation and usage.

bioinformatics↗

Selection analysis identifies unusual clustered mutational changes in Omicron lineage BA.1 that likely impact Spike function

Among the 30 non-synonymous nucleotide substitutions in the Omicron S-gene are 13 that have only rarely been seen in other SARS-CoV-2 sequences. These mutations cluster within three functionally important regions of the S-gene at sites that will likely impact (i) interactions between subunits of the Spike trimer and the predisposition of subunits to shift from down to up configurations, (ii) interactions of Spike with ACE2 receptors, and (iii) the priming of Spike for membrane fusion. We show here that, based on both the rarity of these 13 mutations in intrapatient sequencing reads and patterns of selection at the codon sites where the mutations occur in SARS-CoV-2 and related sarbecoviruses, prior to the emergence of Omicron the mutations would have been predicted to decrease the fitness of any genomes within which they occurred. We further propose that the mutations in each of the three clusters therefore cooperatively interact to both mitigate their individual fitness costs, and adaptively alter the function of Spike. Given the evident epidemic growth advantages of Omicron over all previously known SARS-CoV-2 lineages, it is crucial to determine both how such complex and highly adaptive mutation constellations were assembled within the Omicron S-gene, and why, despite unprecedented global genomic surveillance efforts, the early stages of this assembly process went completely undetected.

evolutionary biology↗