bioRxiv Science⌕ Search

Biology subjects

Ben Nun, N.

Publications and source records attributed to Ben Nun, N..

5 recordsLinked to original sources

Inferring viral proteins that act as public goods during coinfection

Interactions among individuals in structured populations can alter fitness effects of mutations and reshape evolutionary processes. In many systems, including bacteria, yeast, and viruses, such interactions often result in public goods: gene products that are costly to produce yet exploitable by others. During viral coinfection of the same cell, gene products from one genome may complement deleterious mutations in another, allowing defective genomes to persist. Yet it remains difficult to infer which proteins are shareable from population sequencing data, because mutation, selection, drift, and complementation are intertwined. Here, we developed a quantitative framework to infer protein-specific public goods in the RNA bacteriophage MS2, which encodes only four proteins. We analyzed experimental evolution data generated under two multiplicity-of-infection (MOI) regimes: low MOI, where coinfection is rare, and high MOI, where coinfection is common. We first compared empirical mutation patterns between regimes and then applied a Wright-Fisher model combined with simulation-based Bayesian inference using neural posterior estimation. In a two-stage strategy, gene-specific fitness effects were inferred from low-MOI data and subsequently used to estimate protein sharing under high-MOI conditions. Across two statistical inference frameworks, lysis emerged as the strongest public-good candidate, replicase and coat showed an intermediate signal, and maturation showed the weakest evidence for sharing. Together, our results show that viral proteins differ markedly in their propensity to act as public goods. More broadly, they illustrate how coinfection can generate density-dependent selection, a general feature of social evolution that may shape evolutionary dynamics.

evolutionary biology↗

Collective Posterior Inference from Highly Variable Empirical Replicates

High-throughput experimental platforms now routinely generate data from dozens or hundreds of independent observations. Simulation-based inference (SBI) offers a powerful framework for estimating model parameters from such complex datasets, but standard methods struggle to scale to the noisy multiple-replicates regime without incurring prohibitive computational costs or careful hyperparameter tuning. Here, we introduce a new method for fast and robust collective posterior inference from multiple independent replicates using a robust product-of-experts aggregation scheme that automatically mitigates the influence of outliers. Evaluating it on synthetic and empirical evolutionary datasets, we find it achieves state-of-the-art estimation accuracy and computational efficiency, including inference from noisy observations. Our method is compatible with any SBI framework, providing a scalable, plug-and-play solution for inference from noisy multiple-replicate datasets.

bioinformatics↗

Segmental copy number amplifications are stable in the absence of selection

Copy number variants (CNVs) are DNA duplications and deletions that cause genetic variation, underlying rapid adaptive evolution. CNVs often confer selective advantages, but can also incur fitness costs. Evolution of Saccharomyces cerevisiae in nutrient-limited chemostats recurrently selects for amplifications of nutrient transporter genes. However, their fate upon return to a non-selective environment remains unknown. To investigate CNV fitness and stability upon removing the original selection pressure, we studied 15 CNV lineages (11 segmental, 4 whole-chromosomal amplifications) selected in nitrogen-limited chemostats. CNV stability was monitored using fluorescent reporters during propagation in nutrient-rich batch cultures for 110-220 generations. All aneuploid lineages showed rapid CNV loss and reversion to a single-copy genotype, whereas segmental amplifications were remarkably stable- one of 11 strains reverted. Pairwise fitness competitions in rich media revealed strong fitness defects associated solely with CNVs that reverted; reversion led to increased fitness. Using simulation-based inference to estimate reversion rates and fitness effects, we determined negative selection as the primary driver of CNV loss. Whole-genome sequencing revealed that reversion of aneuploids and a segmental amplification left no evidence of prior CNV existence, rendering revertant genomes indistinguishable from the single-copy ancestor. Detailed characterization of a partial revertant identified chromosomal translocation, suggesting that extant CNVs can undergo structural diversification. Our findings provide novel evidence that most segmental CNVs adapted to nitrogen limitation are stable upon removal of selection, but costly gene amplifications are readily reversible. Together, these highlight the importance of CNVs in both long-term genome evolution and rapid, reversible adaptation to transient selection.

evolutionary biology↗

DNA replication errors are a major source of adaptive gene amplification

Copy number variants (CNVs)--gains and losses of genomic sequences--are an important source of genetic variation underlying rapid adaptation and genome evolution. However, despite their central role in evolution little is known about the factors that contribute to the structure, size, formation rate, and fitness effects of adaptive CNVs. Local genomic sequences are likely to be an important determinant of these properties. Whereas it is known that point mutation rates vary with genomic location and local DNA sequence features, the role of genome architecture in the formation, selection, and the resulting evolutionary dynamics of CNVs is poorly understood. Previously, we have found that the GAP1 gene in Saccharomyces cerevisiae undergoes frequent and repeated amplification and selection under long-term experimental evolution in glutamine-limiting conditions. The GAP1 gene has a unique genomic architecture consisting of two flanking long terminal repeats (LTRs) and a proximate origin of DNA replication (autonomously replicating sequence, ARS), which are likely to promote rapid GAP1 CNV formation. To test the role of these genomic elements on CNV-mediated adaptive evolution, we performed experimental evolution in glutamine-limited chemostats using engineered strains lacking either the adjacent LTRs, ARS, or all elements. Using a CNV reporter system and neural network simulation-based inference (nnSBI) we quantified the formation rate and fitness effect of CNVs for each strain. We find that although GAP1 CNVs repeatedly form and sweep to high frequency in strains with modified genome architecture, removal of local DNA elements significantly impacts the rate and fitness effect of CNVs and the rate of adaptation. We performed genome sequence analysis to define the molecular mechanisms of CNV formation for 177 CNV lineages. We find that across all four strain backgrounds, between 26% and 80% of all GAP1 CNVs are mediated by Origin Dependent Inverted Repeat Amplification (ODIRA) which results from template switching between the leading and lagging strand during DNA synthesis. In the absence of the local ARS, a distal ARS can mediate CNV formation via ODIRA. In the absence of local LTRs, homologous recombination mechanisms still mediate gene amplification following de novo insertion of retrotransposon elements at the locus. Our study demonstrates the remarkable plasticity of the genome and reveals that template switching during DNA replication is a frequent source of adaptive CNVs.

evolutionary biology↗

Mutation rate, selection, and epistasis inferred from RNA virus haplotypes via neural posterior estimation

RNA viruses are particularly notorious for their high levels of genetic diversity, which is generated through the forces of mutation and natural selection. However, disentangling these two forces is a considerable challenge, and this may lead to widely divergent estimates of viral mutation rates, as well as difficulties in inferring fitness effects of mutations. Here, we develop, test, and apply an approach aimed at inferring the mutation rate and key parameters that govern natural selection, from haplotype sequences covering full length genomes of an evolving virus population. Our approach employs neural posterior estimation, a computational technique that applies simulation-based inference with neural networks to jointly infer multiple model parameters. We first tested our approach on synthetic data simulated using different mutation rates and selection parameters while accounting for sequencing errors. Reassuringly, the inferred parameter estimates were accurate and unbiased. We then applied our approach to haplotype sequencing data from a serial-passaging experiment with the MS2 bacteriophage. We estimated that the mutation rate of this phage is around 0.2 mutations per genome per replication cycle (95% highest density interval: 0.051-0.56). We validated this finding with two different approaches based on single-locus models that gave similar estimates but with much broader posterior distributions. Furthermore, we found evidence for reciprocal sign epistasis between four strongly beneficial mutations that all reside in an RNA stem-loop that controls the expression of the viral lysis protein, responsible for lysing host cells and viral egress. We surmise that there is a fine balance between over and under-expression of lysis that leads to this pattern of epistasis. To summarize, we have developed an approach for joint inference of the mutation rate and selection parameters from full haplotype data with sequencing errors, and used it to reveal features governing MS2 evolution.

evolutionary biology↗