bioRxiv ScienceSearch

Biology subjects

Ren Sun

Publications and source records attributed to Ren Sun.

6 recordsLinked to original sources

Quantifying the evolutionary potential and constraints of a drug-targeted viral protein

RNA viruses are notorious for their ability to evolve rapidly under selection in novel environments. It is known that the high mutation rate of RNA viruses can generate huge genetic diversity to facilitate viral adaptation. However, less attention has been paid to the underlying fitness landscape that represents the selection forces on viral genomes. Here we systematically quantified the distribution of fitness effects (DFE) of about 1,600 single amino acid substitutions in the drug-targeted region of NS5A protein of Hepatitis C Virus (HCV). We found that the majority of non-synonymous substitutions incur large fitness costs, suggesting that NS5A protein is highly optimized in natural conditions. We characterized the adaptive potential of HCV by subjecting the mutant viruses to selection by the antiviral drug Daclatasvir. Both the selection coefficient and the number of beneficial mutations are found to increase with the level of environmental stress, which is modulated by the concentration of Daclatasvir. The changes in the spectrum of beneficial mutations in NS5A protein can be explained by a pharmacodynamics model describing viral fitness as a function of drug concentration. We test theoretical predictions regarding the distribution of beneficial fitness effects of mutations. We also interpret the data in the context of Fishers Geometric Model and find an increased distance to optimum as a function of environmental stress. Finally, we show that replication fitness of viruses is correlated with the pattern of sequence conservation in nature and viral evolution is constrained by the need to maintain protein stability.

Evolutionary Biology

Comparative analysis of protein evolution and RNA structural changes in the genome of pre-epidemic and epidemic Zika virus

Zika virus (ZIKV) infection is associated with microcephaly, neurological disorders and poor pregnancy outcome1-3 and no vaccine is available. Although ZIKV was first discovered in 1947, the exact mechanism of virus replication and pathogenesis still remains unknown. Recent outbreaks of Zika virus in the Americas clearly suggest a better adaptation of viral strains to human host. Understanding the conserved and adaptive features in the evolution of ZIKV genome will reveal the molecular mechanism of virus replication and host adaptation. Here, we show comprehensive analysis of protein evolution and changes in RNA secondary structures of ZIKV strains including the current 2015-16 outbreak. To identify the constraints on ZIKV evolution, selection pressure at individual codons, immune epitopes, co-evolving sites, and RNA structures were analyzed. The proteome of current 2015/16 epidemic ZIKV strains of Asian genotype is found to be genetically conserved due to genome-wide negative selection on codons, with limited positive selection. Predicted RNA structures at the 5 and 3 ends of ZIKV strains reveal substantial changes such as an additional stem loop which makes it similar to that of Yellow Fever Virus. Concisely, the targeted changes at both the amino acid and the RNA levels contribute to the better adaptation of ZIKV strains to human host with an enhanced neurotropism.

Microbiology

Adaptation in protein fitness landscapes is facilitated by indirect paths

The structure of fitness landscapes is critical for understanding adaptive protein evolution (e.g. antimicrobial resistance, affinity maturation, etc.). Due to limited throughput in fitness measurements, previous empirical studies on fitness landscapes were confined to either the neighborhood around the wild type sequence, involving mostly single and double mutants, or a combinatorially complete subgraph involving only two amino acids at each site. In reality, however, the dimensionality of protein sequence space is higher (20L, L being the length of the relevant sequence) and there may be higher-order interactions among more than two sites. To study how these features impact the course of protein evolution, we experimentally characterized the fitness landscape of four sites in the IgG-binding domain of protein G, containing 204 = 160,000 variants. We found that the fitness landscape was rugged and direct paths of adaptation were often constrained by pairwise epistasis. However, while direct paths were blocked by reciprocal sign epistasis, we found systematic evidence that such evolutionary traps could be circumvented by \"extra-dimensional bypass\". Extra dimensions in sequence space - with a different amino acid at the site of interest or an additional interacting site - open up indirect paths of adaptation via gain and subsequent loss of mutations. These indirect paths alleviate the constraint on reaching high fitness genotypes via selectively accessible trajectories, suggesting that the heretofore neglected dimensions of sequence space may completely change our views on how proteins evolve.

Evolutionary Biology

Long single-molecule reads can resolve the complexity of the Influenza virus composed of rare, closely related mutant variants

As a result of a high rate of mutations and recombination events, an RNA-virus exists as a heterogeneous \"swarm\". The ability of next-generation sequencing to produce massive quantities of genomic data inexpensively has allowed virologists to study the structure of viral populations from an infected host at an unprecedented resolution. However, high similarity and low frequency of the viral variants impose a huge challenge to assembly of individual full-length genomes. The long read length offered by a single-molecule sequencing technologies allows each mutant variant to be sequenced in a single pass. However, high error rate limits the ability to reconstruct heterogeneous viral population composed of rare, related mutant variants. In this paper, we present 2SNV, a method able to tolerate the high error-rate of the single-molecule protocol and reconstruct mutant variants. The proposed protocol is able to eliminate sequencing errors and reconstruct closely related viral mutant variants. 2SNV uses linkage between single nucleotide variations to efficiently distinguish them from read errors. To benchmark the sensitivity of 2SNV, we performed a single-molecule sequencing experiment on a sample containing a titrated level of known viral mutant variants.\n\nOur method is able to accurately reconstruct clone with frequency of 0.2% and distinguish clones that differed in only two nucleotides distantly located on the genome. 2SNV outperforms existing methods for full-length viral mutant reconstruction. With a high sensitivity and accuracy, 2SNV is anticipated to facilitate not only viral quasispecies reconstruction, but also other biological questions that require detection of rare haplotypes such as genetic diversity in cancer cell population, and monitoring B-cell and T-cell receptor repertoire. The open source implementation of 2SNV is freely available for download at http://alan.cs.gsu.edu/NGS/?q=content/2snv

Bioinformatics

Rational design and adaptive management of combination therapies for Hepatitis C virus infection

Recent discoveries of direct acting antivirals against Hepatitis C virus (HCV) have raised hopes of effective treatment via combination therapies. Yet rapid evolution and high diversity of HCV populations, combined with the reality of suboptimal treatment adherence, make drug resistance a clinical and public health concern. We develop a general model incorporating viral dynamics and pharmacokinetics/pharmacodynamics to assess how suboptimal adherence affects resistance development and clinical outcomes. We derive design principles and adaptive treatment strategies, identifying a high-risk period when missing doses is particularly risky for de novo resistance, and quantifying the number of additional doses needed to compensate when doses are missed. Using data from large-scale resistance assays, we demonstrate that the risk of resistance can be reduced substantially by applying these principles to a combination therapy of daclatasvir and asunaprevir. By providing a mechanistic framework to link patient characteristics to the risk of resistance, these findings show the potential of rational treatment design.

Evolutionary Biology

High-throughput functional annotation of influenza A virus genome at single-nucleotide resolution

A novel genome-wide genetics platform is presented in this study, which permits functional interrogation of all point mutations across a viral genome in parallel. Here we generated the first fitness profile of individual point mutations across the influenza virus genome. Critical residues on the viral genome were systematically identified, which provided a collection of subdomain data informative for structure-function studies and for effective rational drug and vaccine design. Our data was consistent with known, well-characterized structural features. In addition, we have achieved a validation rate of 68% for severely attenuated mutations and 94% for neutral mutations. The approach described in this study is applicable to other viral or microbial genomes where a means of genetic manipulation is available.

Systems Biology