bioRxiv Science⌕ Search

Biology subjects

Rodrigues, M. F.

Publications and source records attributed to Rodrigues, M. F..

5 recordsLinked to original sources

Shared evolutionary processes shape landscapes of genomic variation in the great apes

For at least the past five decades population genetics, as a field, has worked to describe the precise balance of forces that shape patterns of variation in genomes. The problem is challenging because modelling the interactions between evolutionary processes is difficult, and different processes can impact genetic variation in similar ways. In this paper, we describe how diversity and divergence between closely related species change with time, using correlations between landscapes of genetic variation as a tool to understand the interplay between evolutionary processes. We find strong correlations between landscapes of diversity and divergence in a well sampled set of great ape genomes, and explore how various processes such as incomplete lineage sorting, mutation rate variation, GC-biased gene conversion and selection contribute to these correlations. Through highly realistic, chromosome-scale, forward-in-time simulations we show that the landscapes of diversity and divergence in the great apes are too well correlated to be explained via strictly neutral processes alone. Our best fitting simulation includes both deleterious and beneficial mutations in functional portions of the genome, in which 9% of fixations within those regions is driven by positive selection. This study provides a framework for modelling genetic variation in closely related species, an approach which can shed light on the complex balance of forces that have shaped genetic variation.

evolutionary biology↗

Expanding the stdpopsim species catalog, and lessons learned forrealistic genome simulations

Simulation is a key tool in population genetics for both methods development and empirical research, but producing simulations that recapitulate the main features of genomic data sets remains a major obstacle. Today, more realistic simulations are possible thanks to large increases in the quantity and quality of available genetic data, and to the sophistication of inference and simulation software. However, implementing these simulations still requires substantial time and specialized knowledge. These challenges are especially pronounced for simulating genomes for species that are not well-studied, since it is not always clear what information is required to produce simulations with a level of realism sufficient to confidently answer a given question. The community-developed framework stdpopsim seeks to lower this barrier by facilitating the simulation of complex population genetic models using up-to-date information. The initial version of stdpopsim focused on establishing this framework using six well-characterized model species (Adrion et al., 2020). Here, we report on major improvements made in the new release of stdpopsim (version 0.2), which includes a significant expansion of the species catalog and substantial additions to simulation capabilities. Features added to improve the realism of the simulated genomes include non-crossover recombination and provision of species-specific genomic annotations. Through community-driven efforts, we expanded the number of species in the catalog more than three-fold and broadened coverage across the tree of life. During the process of expanding the catalog, we have identified common sticking points and developed best practices for setting up genome-scale simulations. We describe the input data required for generating a realistic simulation, suggest good practices for obtaining the relevant information from the literature, and discuss common pitfalls and major considerations. These improvements to stdpopsim aim to further promote the use of realistic whole-genome population genetic simulations, especially in non-model organisms, making them available, transparent, and accessible to everyone.

bioinformatics↗

The origin and evolution of loqs2: a gene encoding an antiviral dsRNA binding protein in Aedes mosquitoes

Mosquito borne viruses such as dengue, Zika, yellow fever and Chikungunya cause millions of infections every year. These viruses are mostly transmitted by two urban-adapted mosquito species, Aedes aegypti and Aedes albopictus, that appear to be more permissive to arbovirus infections compared to closely related species, although mechanistic understanding remains unknown. Aedes mosquitoes may have evolved specialized antiviral mechanisms that potentially contribute to the low impact of viral infection. Recently, we reported the identification of an Aedes specific double-stranded RNA binding protein (dsRBP), named Loqs2, that is involved in the control of infection by dengue and Zika viruses in Ae. aegypti. Loqs2 interacts with two important co-factors of the RNA interference (RNAi) pathway, Loquacious (Loqs) and R2D2, and seems to be a strong regulator of the antiviral defense. However, the origin and evolution of loqs2 remains unclear. Here, we describe that loqs2 likely originated from two independent duplications of the first dsRNA binding domain (dsRBD) of loquacious that occurred before the radiation of the Aedes Stegomyia subgenus. After its origin, our analyses suggest that loqs2 evolved by relaxed positive selection towards neofunctionalization. In fact, loqs2 is evolving at a faster pace compared to other RNAi components such as loquacious, r2d2 and Dicer-2 in Aedes mosquitoes. Unlike loquacious, transcriptomic analysis showed that loqs2 expression is tightly regulated, almost restricted to reproductive tissues in Ae. aegypti and Ae. albopictus. Transgenic mosquitoes engineered to ubiquitously express loqs2 show massive dysregulation of stress response genes and undergo developmental arrest at larval stages. Overall, our results uncover the possible origin and neofunctionalization of a novel antiviral gene, loqs2, in Aedes mosquitoes that ultimately may contribute to their effectiveness as vectors for arboviruses.

evolutionary biology↗

Efficient ancestry and mutation simulation with msprime 1.0

Stochastic simulation is a key tool in population genetics, since the models involved are often analytically intractable and simulation is usually the only way of obtaining ground-truth data to evaluate inferences. Because of this necessity, a large number of specialised simulation programs have been developed, each filling a particular niche, but with largely overlapping functionality and a substantial duplication of effort. Here, we introduce msprime version 1.0, which efficiently implements ancestry and mutation simulations based on the succinct tree sequence data structure and tskit library. We summarise msprimes many features, and show that its performance is excellent, often many times faster and more memory efficient than specialised alternatives. These high-performance features have been thoroughly tested and validated, and built using a collaborative, open source development model, which reduces duplication of effort and promotes software quality via community engagement.

genetics↗

Natural selection and parallel clinal and seasonal changes in Drosophila melanogaster

Spatial and seasonal variation in the environment are ubiquitous. Environmental heterogeneity can affect natural populations and lead to covariation between environment and allele frequencies. Drosophila melanogaster is known to harbor polymorphisms that change both with latitude and seasons. Identifying the role of selection in driving these changes is not trivial, because non-adaptive processes can cause similar patterns. Given the environment changes in similar ways across seasons and along the latitudinal gradient, one promising approach may be to look for parallelism between clinal and seasonal change. Here, we test whether there is a genome-wide correlation between clinal and seasonal change, and whether the pattern is consistent with selection. Allele frequency estimates were obtained from pooled samples from seven different locations along the east coast of the US, and across seasons within Pennsylvania. We show that there is a genome-wide correlation between clinal and seasonal variation, which cannot be explained by linked selection alone. This pattern is stronger in genomic regions with higher functional content, consistent with natural selection. We derive a way to biologically interpret these correlations and show that around 3.7% of the common, autosomal variants could be under parallel seasonal and spatial selection. Our results highlight the contribution of natural selection in driving fluctuations in allele frequencies in natural fly populations and point to a shared genomic basis to climate adaptation which happens over space and time in D. melanogaster.

evolutionary biology↗