bioRxiv ScienceSearch

Biology subjects

David L Robertson

Publications and source records attributed to David L Robertson.

3 recordsLinked to original sources

Identification of important amino acid replacements in the 2013-2016 Ebola virus outbreak

The phylogenetic relationships of Zaire ebolavirus have been intensively analysed over the course of the 2013-2016 outbreak. However, there has been limited consideration of the functional impact of this variation. Here we describe an analysis of the available sequence data in the context of protein structure and phylogenetic history. Amino acid replacements are rare and predicted to have minor effects on protein stability. Synonymous mutations greatly outnumber nonsynonymous mutations, and most of the latter fall into unstructured intrinsically disordered regions, indicating that purifying selection is the dominant mode of selective pressure. However, one replacement, occurring early in the outbreak in Gueckedou in Guinea on 31st March 2014 (alanine to valine at position 82 in the GP protein), is close to the site where the virus binds to the host receptor NPC1 and is located in the phylogenetic tree at the origin of the major B lineage of the outbreak. The functional and evolutionary evidence indicates this A82V change likely has consequences for EBOV's host specificity and hence adaptation to humans.

Microbiology

Ebola virus is evolving but not changing: no evidence for functional change in EBOV from 1976 to the 2014 outbreak

The Ebola epidemic is having a devastating impact in West Africa. Sequencing of Ebola viruses from infected individuals has revealed extensive genetic variation, leading to speculation that the virus may be adapting to the human host and accounting for the scale of the 2014 outbreak. We show that so far there is no evidence for adaptation of EBOV to humans. We analyze the putatively functional changes associated with the current and previous Ebola outbreaks, and find no significant molecular changes. Observed amino acid replacements have minimal effect on protein structure, being neither stabilizing nor destabilizing. Replacements are not found in regions of the proteins associated with known functions and tend to occur in disordered regions. This observation indicates that the difference between the current and previous outbreaks is not due to the observed evolutionary change of the virus. Instead, epidemiological factors must be responsible for the unprecedented spread of EBOV.

Evolutionary Biology

Alignment by numbers: sequence assembly using compressed numerical representations

MotivationDNA sequencing instruments are enabling genomic analyses of unprecedented scope and scale, widening the gap between our abilities to generate and interpret sequence data. Established methods for computational sequence analysis generally use nucleotide-level resolution of sequences, and while such approaches can be very accurate, increasingly ambitious and data-intensive analyses are rendering them impractical for applications such as genome and metagenome assembly. Comparable analytical challenges are encountered in other data-intensive fields involving sequential data, such as signal processing, in which dimensionality reduction methods are routinely used to reduce the computational burden of analyses. We therefore seek to address the question of whether it is possible to improve the efficiency of sequence alignment by applying dimensionality reduction methods to numerically represented nucleotide sequences.\n\nResultsTo explore the applicability of signal transformation and dimensionality reduction methods to sequence assembly, we implemented a short read aligner and evaluated its performance against simulated high diversity viral sequences alongside four existing aligners. Using our sequence transformation and feature selection approach, alignment time was reduced by up to 14-fold compared to uncompressed sequences and without reducing alignment accuracy. Despite using highly compressed sequence transformations, our implementation yielded alignments of similar overall accuracy to existing aligners, outperforming all other tools tested at high levels of sequence variation. Our approach was also applied to the de novo assembly of a simulated diverse viral population. Our results demonstrate that full sequence resolution is not a prerequisite of accurate sequence alignment and that analytical performance can be retained and even enhanced through appropriate dimensionality reduction of sequences.

Bioinformatics