bioRxiv ScienceSearch

Biology subjects

Koonin, E. V.

Publications and source records attributed to Koonin, E. V..

14 recordsLinked to original sources

Origins and Evolution of the Global RNA Virome

Viruses with RNA genomes dominate the eukaryotic virome, reaching enormous diversity in animals and plants. The recent advances of metaviromics prompted us to perform a detailed phylogenomic reconstruction of the evolution of the dramatically expanded global RNA virome. The only universal gene among RNA viruses is the RNA-dependent RNA polymerase (RdRp). We developed an iterative computational procedure that alternates the RdRp phylogenetic tree construction with refinement of the underlying multiple sequence alignments. The resulting tree encompasses 4,617 RNA virus RdRps and consists of 5 major branches, 2 of which include positive-sense RNA viruses, 1 is a mix of positive-sense (+) RNA and double-stranded (ds) RNA viruses, and 2 consist of dsRNA and negative-sense (-) RNA viruses, respectively. This tree topology implies that dsRNA viruses evolved from +RNA viruses on at least two independent occasions, whereas -RNA viruses evolved from dsRNA viruses. Reconstruction of RNA virus evolution using the RdRp tree as the scaffold suggests that the last common ancestors of the major branches of +RNA viruses encoded only the RdRp and a single jelly-roll capsid protein. Subsequent evolution involved independent capture of additional genes, particularly, those encoding distinct RNA helicases, enabling replication of larger RNA genomes and facilitating virus genome expression and virus-host interactions. Phylogenomic analysis reveals extensive gene module exchange among diverse viruses and horizontal virus transfer between distantly related hosts. Although the network of evolutionary relationships within the RNA virome is bound to further expand, the present results call for a thorough reevaluation of the RNA virus taxonomy.\n\nIMPORTANCEThe majority of the diverse viruses infecting eukaryotes have RNA genomes, including numerous human, animal, and plant pathogens. Recent advances of metagenomics have led to the discovery of many new groups of RNA viruses in a wide range of hosts. These findings enable a far more complete reconstruction of the evolution of RNA viruses than what was attainable previously. This reconstruction reveals the relationships between different Baltimore Classes of viruses and indicates extensive transfer of viruses between distantly related hosts, such as plants and animals. These results call for a major revision of the existing taxonomy of RNA viruses.

microbiology

Stability of host-parasite systems: you must differ to coevolve

BackgroundGenetic parasites are ubiquitous satellites of cellular life forms most of which host a variety of mobile genetic elements including transposons, plasmids and viruses. Theoretical considerations and computer simulations suggest that emergence of genetic parasites is intrinsic to evolving replicator systems.\n\nResultsUsing methods of bifurcation analysis, we investigated the stability of simple models of replicator-parasite coevolution in a well-mixed environment. It is shown that the simplest imaginable system of this type, one in which the parasite evolves during the replication of the host genome through a minimal mutation that renders the genome of the emerging parasite incapable of producing the replicase but able to recognize and recruit it for its own replication, is unstable. In this model, there are only either trivial or \"semi-trivial\", parasite-free equilibria: an inefficient parasite is outcompeted by the host and dies off whereas an efficient one pushes the host out of existence, which leads to the collapse of the entire system. We show that stable host-parasite coevolution (a non-trivial equilibrium) is possible in a modified model where the parasite is qualitatively distinct from the host replicator in that the replication of the parasite depends solely on the availability of the host but not on the carrying capacity of the environment.\n\nConclusionsWe analytically determine the conditions for stable host-parasite coevolution in simple mathematical models and find that a parasite that initially evolves from the host through the loss of the ability to replicate autonomously must be substantially derived for a stable host-parasite coevolution regime to be established.

evolutionary biology

On the feasibility of saltational evolution

One of the key tenets of Darwins theory that was inherited by the Modern Synthesis of evolutionary biology is gradualism, that is, the notion that evolution proceeds gradually, via accumulation of \"infinitesimally small\" heritable changes 1,2. However, some of the most consequential evolutionary changes, such as, for example, the emergence of major taxa, seem to occur abruptly rather than gradually, as captured in the concepts of punctuated equilibrium 3,4 and evolutionary transitions 5,6. We examine a mathematical model of an evolutionary process on a rugged fitness landscape 7,8 and obtain analytic solutions for the probability of multi-mutational leaps, that is, several mutations occurring simultaneously, within a single generation in one genome, and being fixed all together in the evolving population. The results indicate that, for typical, empirically observed combinations of the parameters of the evolutionary process, namely, effective population size, mutation rate, and distribution of selection coefficients of mutations, the probability of a multi-mutational leap is low, and accordingly, their contribution to the evolutionary process is minor at best. However, such leaps could become an important factor of evolution in situations of population bottlenecks and elevated mutation rates, such as stress-induced mutagenesis in microbes or tumor progression, as well as major evolutionary transitions and evolution of primordial replicators.

evolutionary biology

Genome plasticity, a key factor of evolution in prokaryotes

In prokaryotic genomes, the number of genes that belong to distinct functional classes shows apparent universal scaling with the total number of genes [1-5] (Fig. 1). This scaling can be approximated with a power law, where the scaling power can be sublinear, near-linear or super-linear. Scaling laws are robust under various statistical tests [4], across different databases and for different gene classifications [1-5]. Several models aimed at explaining the observed scaling laws have been proposed, primarily, based on the specifics of the respective biological functions [1, 5-8]. However, a coherent theory to explain the emergence of scaling within the framework of population genetics is lacking. We employ a simple mathematical model for prokaryotic genome evolution [9] which, together with the analysis of 34 clusters of closely related microbial genomes [10], allows us to identify the underlying forces that dictate genome content evolution. In addition to the scaling of the number of genes in different functional classes, we explore gene contents divergence to characterize the evolutionary processes acting upon genomes [11]. We find that evolution of the gene content is dominated by two factors that are specific to a functional class, namely, selection landscape and genome plasticity. Selection landscape quantifies the fitness cost that is associated with deletion of a gene in a given functional class or the advantage of successful incorporation of an additional gene. Genome plasticity, that can be considered a measure of evolvability, reflects both the availability of the genes of a given functional class in the external gene pool that is accessible to the evolving microbial population, and the ability of microbial genomes to accommodate these genes. The selection landscape determines the gene loss rate, and genome plasticity is the principal determinant of the gene gain rate.\n\nO_FIG O_LINKSMALLFIG WIDTH=197 HEIGHT=200 SRC=\"FIGDIR/small/357400_fig1.gif\" ALT=\"Figure 1\">\nView larger version (39K):\norg.highwire.dtl.DTLVardef@6df3e2org.highwire.dtl.DTLVardef@a69e8dorg.highwire.dtl.DTLVardef@f36a80org.highwire.dtl.DTLVardef@d519c9_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Scaling laws for all functional classes of the COGs. The number of genes in a given COG category is plotted against the total number of genes. Each point represents one genome from the analyzed set of 1490 genomes. The scaling is fitted to a power law which is indicated by a solid red line. The fitted scaling exponent is indicated in parentheses.\n\nC_FIG

microbiology

Adomaviruses: an emerging virus family provides insights into DNA virus evolution

Adenoviruses, papillomaviruses, and polyomaviruses are collectively known as small DNA tumor viruses. Although it has long been recognized that small DNA tumor virus oncoproteins and capsid proteins show a variety of structural and functional similarities, it is unclear whether these similarities reflect descent from a common ancestor, convergent evolution, horizontal gene transfer among virus lineages, or acquisition of genes from host cells. Here, we report the discovery of a dozen new members of an emerging virus family, the Adomaviridae, that unite a papillomavirus/polyomavirus-like replicase gene with an adenovirus-like virion maturational protease. Adomaviruses were initially discovered in a lethal disease outbreak among endangered Japanese eels. New adomavirus genomes were found in additional commercially important fish species, such as tilapia, as well as in reptiles. The search for adomavirus sequences also revealed an additional candidate virus family, which we refer to as xenomaviruses, in mollusk datasets. Analysis of native adomavirus virions and expression of recombinant proteins showed that the virion structural proteins of adomaviruses are homologous to those of both adenoviruses and another emerging animal virus family called adintoviruses. The results pave the way toward development of vaccines against adomaviruses and suggest a framework that ties small DNA tumor viruses into a shared evolutionary history. Author SummaryIn contrast to cellular organisms, viruses do not encode any universally conserved genes. Even within a given family of viruses, the amino acid sequences encoded by homologous genes can diverge to the point of unrecognizability. Although members of an emerging virus family, the Adomaviridae, encode replicative DNA helicase proteins that are recognizably similar to those of polyomaviruses and papillomaviruses, the functions of other adomavirus genes have been difficult to identify. Using a combination of laboratory and bioinformatic approaches, we identify the adomavirus virion structural proteins. The results link adomavirus virion protein operons to those of other midsize non-enveloped DNA viruses, including adenoviruses and adintoviruses.

microbiology

Criticality in Tumor Evolution and Clinical Outcome

How mutation and selection determine the fitness landscape of tumors and hence clinical outcome is an open fundamental question in cancer biology, crucial for the assessment of therapeutic strategies and resistance to treatment. Here we explore the mutation-selection phase-diagram of 6721 primary tumors representing 23 cancer types, by quantifying the overall somatic point mutation load (ML) and selection (dN/dS) in the entire proteome of each tumor. We show that ML strongly correlates with patient survival, revealing two opposing regimes around a critical point. In low ML cancers, high number of mutations indicates poor prognosis, whereas high ML cancers show the opposite trend, due to mutational meltdown. Although the majority of cancers evolve near neutrality, deviations are observed at extreme MLs. Cancers with the highest ML evolve under purifying selection, whereas those with the lowest ML show signatures of positive selection, demonstrating how selection affects cancer fitness. Moreover, different cancers occupy specific positions on the ML-dN/dS plane, revealing a diversity of evolutionary trajectories. These results support and expand the theory of tumor evolution and its non-linear effects on survival.\n\nSignificance StatementIt remains an open fundamental question how mutation and selection co-determine the course of cancer evolution. We construct a selection-mutation phase diagram, using tumor mutation load and selection strength as key variables, and assess their association with clinical outcome. We demonstrate the existence of a biphasic evolutionary regime, whereby beyond a critical ML, the fitness of tumors decreases with the number of mutations, while the proteome evolves near neutrality. Deviations from neutrality in extreme ML elucidate how positive and purifying selections maintain tumor fitness. These results empirically corroborate the existence of a critical state in cancer evolution predicted by theory, and have fundamental and likely clinical implications.

cancer biology

Towards comprehensive characterization of CRISPR-linked genes

The CRISPR-Cas systems of bacterial and archaeal adaptive immunity consist of arrays of direct repeats separated by unique spacers and multiple CRISPR-associated (cas) genes encoding proteins that mediate the adaptation, CRISPR RNA maturation and interference stages of the CRISPR response. In addition to the relatively small set of core cas genes that are typically present in all representatives of each (sub)type of CRISPR-Cas systems and are essential for the defense function, numerous genes occur in CRISPR-cas loci only sporadically. Some of these have been shown to perform various ancillary roles in CRISPR response whereas the functional relevance of many others, if any, remains obscure. We developed a computational strategy for systematically detecting genes that are likely to be functionally linked to CRISPR-Cas systems. The approach is based on a \"CRISPRicity\" metric that measures the strength of CRISPR association for all protein-coding genes from sequenced bacterial and archaeal genomes. Uncharacterized genes with CRISPRicity values comparable to those of known cas genes are considered candidate CRISPR-ancillary genes, and we describe additional criteria to identify functionally relevant genes in the candidate set. About 80 genes that were not previously reported to be associated with CRISPR-Cas were identified as probable CRISPR-ancillary genes. A substantial majority of these genes reside in type III CRISPR-cas loci which implies exceptional functional versatility of type III systems. Numerous candidate CRISPR-ancillary genes encode integral membrane proteins suggestive of tight membrane connections of type III CRISPR-Cas whereas many other candidates are proteins implicated in various signal transduction pathways. These predictions provide ample material for improving annotation of CRISPR-cas loci and experimental characterization of previously unsuspected aspects of CRISPR-Cas functionality.\n\nSIGNIFICANCEThe CRISPR-Cas systems that mediate adaptive immunity in bacteria and archaea encompass a small set of core cas genes that are essential in a broad range of CRISPR-Cas systems. However, a much greater number of genes only sporadically co-occur with CRISPR-Cas, and for most of these, involvement in CRISPR-Cas functions has not been demonstrated. We developed a computational strategy that provides for systematic identification of CRISPR-linked proteins and prediction of their functional association with CRISPR-Cas systems. About 80 previously undetected, putative CRISPR-accessory proteins were identified. A large fraction of these proteins are predicted to be membrane-associated revealing an unknown side of CRISPR biology.

genomics

The cancer-mutation network and the number and specificity of driver mutations

Cancer genomics has produced extensive information on cancer-associated genes but the number and specificity of cancer driver mutations remains a matter of debate. We constructed a bipartite network in which 7665 tumors from 30 cancer types are connected via shared mutations in 198 previously identified cancer-associated genes. We show that 27% of the tumors can be assigned to statistically supported modules, most of which encompass 1-2 cancer types. The rest of the tumors belong to a diffuse network component suggesting lower gene-specificity of driver mutations. Linear regression of the mutational loads in cancer-associated genes was used to estimate the number of drivers required for the onset of different cancers. The mean number of drivers is ~2, with a range of 1 to 5. Cancers that are associated to modules had more drivers than those from the diffuse network component, suggesting that unidentified and/or interchangeable drivers exist in the latter.

cancer biology

Towards physical principles of biological evolution

Biological systems reach organizational complexity that far exceeds the complexity of any known inanimate objects. Biological entities undoubtedly obey the laws of quantum physics and statistical mechanics. However, is modern physics sufficient to adequately describe, model and explain the evolution of biological complexity? Detailed parallels have been drawn between statistical thermodynamics and the population-genetic theory of biological evolution. Based on these parallels, we outline new perspectives on biological innovation and major transitions in evolution, and introduce a biological equivalent of thermodynamic potential that reflects the innovation propensity of an evolving population. Deep analogies have been suggested to also exist between the properties of biological entities and processes, and those of frustrated states in physics, such as glasses. Such systems are characterized by frustration whereby local state with minimal free energy conflict with the global minimum, resulting in \"emergent phenomena\". We extend such analogies by examining frustration-type phenomena, such as conflicts between different levels of selection, in biological evolution. These frustration effects appear to drive the evolution of biological complexity. We further address evolution in multidimensional fitness landscapes from the point of view of percolation theory and suggest that percolation at level above the critical threshold dictates the tree-like evolution of complex organisms. Taken together, these multiple connections between fundamental processes in physics and biology imply that construction of a meaningful physical theory of biological evolution might not be a futile effort. However, it is unrealistic to expect that such a theory can be created in one scoop; if it ever comes to being, this can only happen through integration of multiple physical models of evolutionary processes. Furthermore, the existing framework of theoretical physics is unlikely to suffice for adequate modeling of the biological level of complexity, and new developments within physics itself are likely to be required.

evolutionary biology

Inevitability of the emergence and persistence of genetic parasites caused by thermodynamic instability of parasite-free states

Genetic parasites, including viruses and mobile genetic elements, are ubiquitous among cellular life forms, and moreover, are the most abundant biological entities on earth that harbor the bulk of the genetic diversity. Here we examine simple thought experiments to demonstrate that both the emergence of parasites in simple replicator systems and their persistence in evolving life forms are inevitable because the putative parasite-free states are thermodynamically unstable.

evolutionary biology

Recruitment of CRISPR-Cas systems by Tn7-like transposons

A survey of bacterial and archaeal genomes shows that many Tn7-like transposons contain minimal type I-F CRISPR-Cas systems that consist of fused cas8f and cas5f, cas7f and cas6f genes, and a short CRISPR array. Additionally, several small groups of Tn7-like transposons encompass similarly truncated type I-B CRISPR-Cas systems. This gene composition of the transposon-associated CRISPR-Cas systems implies that they are competent for pre-crRNA processing yielding mature crRNAs and target binding but not target cleavage that is required for interference. Here we present phylogenetic analysis demonstrating that evolution of the CRISPR-Cas containing transposons included a single, ancestral capture of a type I-F locus and two independent instances of type I-B loci capture. We further show that the transposon-associated CRISPR arrays contain spacers homologous to plasmid and temperate phage sequences, and in some cases, chromosomal sequences adjacent to the transposon. A hypothesis is proposed that the transposon-encoded CRISPR-Cas systems generate displacement (R-loops) in the cognate DNA sites, targeting the transposon to these sites and thus facilitating their spread via plasmids and phages. This scenario fits the \"guns for hire\" concept whereby mobile genetic elements can capture host defense systems and repurpose them for different stages in the life cycle of the element.\n\nImportanceCRISPR-Cas is an adaptive immunity system that protects bacteria and archaea from mobile genetic elements. We present comparative genomic and phylogenetic analysis of degenerate CRISPR-Cas variants associated with distinct families of transposable elements and develop the hypothesis that such repurposed defense systems contribute to the transposable element propagation by facilitating transposition into specific sites. Such recruitment of defense systems by mobile elements supports the \"guns for hire\" concept under which the same enzymatic machineries can be alternately employed for transposon proliferation or host defense.

evolutionary biology

Disentangling The Effects Of Selection And Loss Bias On Gene Dynamics

We combine mathematical modelling of genome evolution with comparative analysis of prokaryotic genomes to estimate the relative contributions of selection and intrinsic loss bias to the evolution of different functional classes of genes and mobile genetic elements (MGE). An exact solution for the dynamics of gene family size was obtained under a linear duplication-transfer-loss model with selection. With the exception of genes involved in information processing, particularly translation, which are maintained by strong selection, the average selection coefficient for most non-parasitic genes is low albeit positive, compatible with the observed positive correlation between genome size and effective population size. Free-living microbes evolve under stronger selection for gene retention than parasites. Different classes of MGE show a broad range of fitness effects, from the nearly neutral transposons to prophages, which are actively eliminated by selection. Genes involved in anti-parasite defense, on average, incur a fitness cost to the host that is at least as high as the cost of plasmids. This cost is probably due to the adverse effects of autoimmunity and curtailment of horizontal gene transfer caused by the defense systems and selfish behavior of some of these systems, such as toxin-antitoxin and restriction-modification modules. Transposons follow a biphasic dynamics, with bursts of gene proliferation followed by decay in the copy number that is quantitatively captured by the model. The horizontal gene transfer to loss ratio, but not the duplication to loss ratio, correlates with genome size, potentially explaining the increased abundance of neutral and costly elements in larger genomes.\n\nSIGNIFICANCEEvolution of microbes is dominated by horizontal gene transfer and the incessant host-parasite arms race that promotes the evolution of diverse anti-parasite defense systems. The evolutionary factors governing these processes are complex and difficult to disentangle but the rapidly growing genome databases provide ample material for testing evolutionary models. Rigorous mathematical modeling of evolutionary processes, combined with computer simulation and comparative genomics, allowed us to elucidate the evolutionary regimes of different classes of microbial genes. Only genes involved in key informational and metabolic pathways are subject to strong selection whereas most of the others are effectively neutral or even burdensome. Mobile genetic elements and defense systems are costly, supporting the understanding that their evolution is governed by the same factors.

evolutionary biology

Novel Abundant Oceanic Viruses of Uncultured Marine Group II Euryarchaeota Identified by Genome-Centric Metagenomics

Marine Group II Euryarchaeota (MGII) are among the most abundant microbes in the oceanic surface waters. So far, however, representatives of MGII have not been cultivated, and no viruses infecting these organisms have been described. Here we present complete genomes for 3 distinct groups of viruses assembled from metagenomic sequence datasets highly enriched for MGII. These novel viruses, which we denote Magroviruses, possess double-stranded DNA genomes of 65 to 100 kilobase in size that encode a structural module characteristic of head-tailed viruses and, unusually for archaeal and bacterial viruses, a nearly complete replication apparatus of apparent archaeal origin. The newly identified Magroviruses are widespread and abundant, and therefore are likely to be major ecological agents.

microbiology

Cas13b is a Type VI-B CRISPR-associated RNA-Guided RNase differentially regulated by accessory proteins Csx27 and Csx28

CRISPR-Cas adaptive immune systems defend microbes against foreign nucleic acids via RNA-guided endonucleases. Using a computational sequence database mining approach, we identify two Class 2 CRISPR-Cas systems (subtype VI-B) that lack Cas1 and Cas2 and encompass a single large effector protein, Cas13b, along with one of two previously uncharacterized associated proteins, Csx27 or Csx28. We establish that these CRISPR-Cas systems can achieve RNA interference when heterologously expressed. Through a combination of biochemical and genetic experiments, we show that Cas13b processes its own CRISPR array with short and long direct repeats, cleaves target RNA, and exhibits collateral RNase activity. Using an E. coli essential gene screen, we demonstrate that Cas13b has a double-sided protospacer-flanking sequence and elucidate RNA secondary structure requirements for targeting. We also find that Csx27 represses, whereas Csx28 enhances, Cas13b-mediated RNA interference. Characterization of these CRISPR systems creates opportunities to develop tools to manipulate and monitor cellular transcripts.

microbiology