bioRxiv ScienceSearch

Biology subjects

Hahn, M. W.

Publications and source records attributed to Hahn, M. W..

14 recordsLinked to original sources

Leveraging evolutionary relationships to improve Anopheles genome assemblies

While new sequencing technologies have lowered financial barriers to whole genome sequencing, resulting assemblies are often fragmented and far from finished. Subsequent improvements towards chromosomal-level status can be achieved by both experimental and computational approaches. Requiring only annotated assemblies and gene orthology data, comparative genomics approaches that aim to capture evolutionary signals to predict scaffold neighbours (adjacencies) offer potentially substantive improvements without the costs associated with experimental scaffolding or re-sequencing. We leverage the combined detection power of three such gene synteny-based methods applied to 21 Anopheles mosquito assemblies with variable contiguity levels to produce consensus sets of scaffold adjacency predictions. Three complementary validations were performed on subsets of assemblies with additional supporting data: six with physical mapping data; 13 with paired-end RNA sequencing (RNAseq) data; and three with new assemblies based on re-scaffolding or incorporating Pacific Biosciences (PacBio) sequencing data. Improved assemblies were built by integrating the consensus adjacency predictions with supporting experimental data, resulting in 20 new reference assemblies with improved contiguities. Combined with physical mapping data for six anophelines, chromosomal positioning of scaffolds improved assembly anchoring by 47% for A. funestus and 38% A. stephensi. Reconciling an A. funestus PacBio assembly with synteny-based and RNAseq-based adjacencies and physical mapping data resulted in a new 81.5% chromosomally mapped reference assembly and cytogenetic photomap. While complementary experimental data are clearly key to achieving high-quality chromosomal-level assemblies, our assessments and validations of gene synteny-based computational methods highlight the utility of applying comparative genomics approaches to improve community genomic resources.

genomics

The perils of intralocus recombination for inferences of molecular convergence

Accurate inferences of convergence require that the appropriate tree topology be used. If there is a mismatch between the tree a trait has evolved along and the tree used for analysis, then false inferences of convergence (\"hemiplasy\") can occur. To avoid problems of hemiplasy when there are high levels of gene tree discordance with the species tree, researchers have begun to construct tree topologies from individual loci. However, due to intralocus recombination even locus-specific trees may contain multiple topologies within them. This implies that the use of individual tree topologies discordant with the species tree can still lead to incorrect inferences about molecular convergence. Here we examine the frequency with which single exons and single protein-coding genes contain multiple underlying tree topologies, in primates and Drosophila, and quantify the effects of hemiplasy when using trees inferred from individual loci. In both clades we find that there are most often multiple diagnosable topologies within single exons and whole genes, with 91% of Drosophila protein-coding genes containing multiple topologies. Because of this underlying topological heterogeneity, even using trees inferred from individual protein-coding genes results in 25% and 38% of substitutions falsely labeled as convergent in primates and Drosophila, respectively. While constructing local trees can reduce the problem of hemiplasy, our results suggest that it will be difficult to completely avoid false inferences of convergence. We conclude by suggesting several ways forward in the analysis of convergent evolution, for both molecular and morphological characters.

evolutionary biology

Quantifying the risk of hemiplasy in phylogenetic inference

Convergent evolution is often inferred when a trait is incongruent with the species tree. However, trait incongruence can also arise from changes that occur on discordant gene trees, a process referred to as hemiplasy. Hemiplasy is rarely taken into account in studies of convergent evolution, despite the fact that phylogenomic studies have revealed rampant discordance. Here, we study the relative probabilities of homoplasy (including convergence and reversal) and hemiplasy for an incongruent trait. We derive expressions for the probabilities of the two events, showing that they depend on many of the same parameters. We find that hemiplasy is as likely-- or more likely--than homoplasy for a wide range of conditions, even when levels of discordance are low. We also present a new method to calculate the ratio of these two probabilities (the \"hemiplasy risk factor\") along the branches of a phylogeny of arbitrary length. Such calculations can be applied to any tree in order to identify when and where incongruent traits may be more likely to be due to hemiplasy than homoplasy.

evolutionary biology

Population genetic tests for the direction and relative timing of introgression

Introgression is a pervasive biological process, and many statistical methods have been developed to infer its presence from genomic data. However, many of the consequences and genomic signatures of introgression remain unexplored from a methodological standpoint. Here, we develop a model for the timing and direction of introgression based on the multispecies network coalescent, and from it suggest new approaches for testing introgression hypotheses. We suggest two new statistics, D1 and D2, which can be used in conjunction with other information to test hypotheses relating to the timing and direction of introgression, respectively. D1 may find use in evaluating cases of homoploid hybrid speciation, while D2 provides a four-taxon test for polarizing introgression. Although analytical expectations for our statistics require a number of assumptions to be met, we show how simulations can be used to test hypotheses about introgression when these assumptions are violated. We apply the D1 statistic to genomic data from the wild yeast Saccharomyces paradoxus, a proposed example of homoploid hybrid speciation, demonstrating its use as a test of this model. These methods provide new and powerful ways to address questions relating to the timing and direction of introgression.

evolutionary biology

Reproductive longevity predicts mutation rates in primates

Mutation rates vary between species across several orders of magnitude, with larger organisms having the highest per-generation mutation rates. Hypotheses for this pattern typically invoke physiological or population-genetic constraints imposed on the molecular machinery preventing mutations1. However, continuing germline cell division in multicellular eukaryotes means that organisms with longer generation times and of larger size will leave more mutations to their offspring simply as a by-product of their increased lifespan2,3. Here, we deeply sequence the genomes of 30 owl monkeys (Aotus nancymaae) from 6 multi-generation pedigrees to demonstrate that paternal age is the major factor determining the number of de novo mutations in this species. We find that owl monkeys have an average mutation rate of 0.81 x 10-8 per site per generation, roughly 32% lower than the estimate in humans. Based on a simple model of reproductive longevity that does not require any changes to the mutational machinery, we show that this is the expected mutation rate in owl monkeys. We further demonstrate that our model predicts species-specific mutation rates in other primates, including study-specific mutation rates in humans based on the average paternal age. Our results suggest that variation in life history traits alone can explain variation in the per-generation mutation rate among primates, and perhaps among a wide range of multicellular organisms.

evolutionary biology

Three new genome assemblies support a rapid radiation in Musa acuminata (wild banana)

Edible bananas result from interspecific hybridization between Musa acuminata and Musa balbisiana, as well as among subspecies in M. acuminata. Four particular M. acuminata subspecies have been proposed as the main contributors of edible bananas, all of which radiated in a short period of time in southeastern Asia. Clarifying the evolution of these lineages at a whole-genome scale is therefore an important step toward understanding the domestication and diversification of this crop. This study reports the de novo genome assembly and gene annotation of a representative genotype from three different subspecies of M. acuminata. These data are combined with the previously published genome of the fourth subspecies to investigate phylogenetic relationships and genome evolution. Analyses of shared and unique gene families reveal that the four subspecies are quite homogenous, with a core genome representing at least 50% of all genes and very few M. acuminata species-specific gene families. Multiple alignments indicate high sequence identity between homologous single copy-genes, supporting the close relationships of these lineages. Interestingly, phylogenomic analyses demonstrate high levels of gene tree discordance, due to both incomplete lineage sorting and introgression. This pattern suggests rapid radiation within Musa acuminata subspecies that occurred after the divergence with M. balbisiana. Introgression between M. a. ssp. malaccensis and M. a. ssp. burmannica was detected across a substantial portion of the genome, though multiple approaches to resolve the subspecies tree converged on the same topology. To support future evolutionary and functional analyses, we introduce the PanMusa database, which enables researchers to exploration of individual gene families and trees.

evolutionary biology

Evolutionary inferences about quantitative traits are affected by underlying genealogical discordance

Modern phylogenetic methods used to study how traits evolve often require a single species tree as input, and do not take underlying gene tree discordance into account. Such approaches may lead to errors in phylogenetic inference because of hemiplasy -- the process by which single changes on discordant trees appear to be homoplastic when analyzed on a fixed species tree. Hemiplasy has been shown to affect inferences about discrete traits, but it is still unclear whether complications arise when quantitative traits are analyzed. In order to address this question and to characterize the effect of hemiplasy on traits controlled by a large number of loci, we present a multispecies coalescent model for quantitative traits evolving along a species tree. We demonstrate theoretically and through simulations that hemiplasy decreases the expected covariances in trait values between more closely related species relative to the covariances between more distantly related species. This effect leads to an overestimation of a traits evolutionary rate parameter, to a decrease of the traits phylogenetic signal, and to increased false positive rates in comparative methods such as the phylogenetic ANOVA. We also show that hemiplasy affects discrete, threshold traits that have an underlying continuous liability, leading to false inferences of convergent evolution. The number of loci controlling a quantitative trait appears to be irrelevant to the trends reported, for all analyses. Our results demonstrate that gene tree discordance and hemiplasy are a problem for all types of traits, across a wide range of methods. Our analyses also point to the conditions under which hemiplasy is most likely to be a factor, and suggest future approaches that may mitigate its effects.

evolutionary biology

Speciation genes are more likely to have discordant gene trees

Speciation genes are responsible for reproductive isolation between species. By directly participating in the process of speciation, the genealogies of isolating loci have been thought to more faithfully represent species trees. The unique properties of speciation genes may provide valuable evolutionary insights and help determine the true history of species divergence. Here, we formally analyze whether genealogies from loci participating in Dobzhansky-Muller (DM) incompatibilities are more likely to be concordant with the species tree under incomplete lineage sorting (ILS). Individual loci differ stochastically from the true history of divergence with a predictable frequency due to ILS, and these expectations--combined with the DM model of intrinsic reproductive isolation from epistatic interactions--can be used to examine the probability of concordance at isolating loci. Contrary to existing verbal models, we find that reproductively isolating loci that follow the DM model are often more likely to have discordant gene trees. These results are dependent on the pattern of isolation observed between three species, the time between speciation events, and the time since the last speciation event. Results supporting a higher probability of discordance are found for both derived-derived and derived-ancestral DM pairs, and regardless of whether incompatibilities are allowed or prohibited from segregating in the same population. Our overall results suggest that DM loci are unlikely to be especially useful for reconstructing species relationships, even in the presence of gene flow between incipient species, and may in fact be positively misleading.

evolutionary biology

Dissecting the basis of novel trait evolution in a radiation with widespread phylogenetic discordance

Phylogenetic analyses of trait evolution can provide insight into the evolutionary processes that initiate and drive phenotypic diversification. However, recent phylogenomic studies have revealed extensive gene tree-species tree discordance, which can lead to incorrect inferences of trait evolution if only a single species tree is used for analysis. This phenomenon--dubbed \"hemiplasy\"--is particularly important to consider during analyses of character evolution in rapidly radiating groups, where discordance is widespread. Here we generate whole-transcriptome data for a phylogenetic analysis of 14 species in the plant genus Jaltomata (the sister clade to Solanum), which has experienced rapid, recent trait evolution, including in fruit and nectar color, and flower size and shape. Consistent with other radiations, we find evidence for rampant gene tree discordance due to incomplete lineage sorting (ILS) and several introgression events among the well-supported subclades. Since both ILS and introgression increase the probability of hemiplasy, we perform several analyses that take discordance into account while identifying genes that might contribute to phenotypic evolution. Despite discordance, the history of fruit color evolution in Jaltomata can be inferred with high confidence, and we find evidence of de novo adaptive evolution at individual genes associated with fruit color variation. In contrast, hemiplasy appears to strongly affect inferences about floral character transitions in Jaltomata, and we identify candidate loci that could arise either from multiple lineage-specific substitutions or standing ancestral polymorphisms. Our analysis provides a generalizable example of how to manage discordance when identifying loci associated with trait evolution in a radiating lineage.

evolutionary biology

Multinucleotide mutations cause false inferences of positive selection

Phylogenetic tests of adaptive evolution, which infer positive selection from an excess of nonsynonymous changes, assume that nucleotide substitutions occur singly and independently. But recent research has shown that multiple errors at adjacent sites often occur in single events during DNA replication. These multinucleotide mutations (MNMs) are overwhelmingly likely to be nonsynonymous. We therefore evaluated whether phylogenetic tests of adaptive evolution, such as the widely used branch-site test, might misinterpret sequence patterns produced by MNMs as false support for positive selection. We explored two genome-wide datasets comprising thousands of coding alignments - one from mammals and one from flies - and found that codons with multiple differences (CMDs) account for virtually all the support for lineage-specific positive selection inferred by the branch-site test. Simulations under genome-wide, empirically derived conditions without positive selection show that realistic rates of MNMs cause a strong and systematic bias in the branch-site and related tests; the bias is sufficient to produce false positive inferences approximately as often as the branch-site test infers positive selection from the empirical data. Our analysis indicates that genes may often be inferred to be under positive selection simply because they stochastically accumulated one or a few MNMs. Because these tests do not reliably distinguish sequence patterns produced by authentic positive selection from those caused by neutral fixation of MNMs, many published inferences of adaptive evolution using these techniques may therefore be artifacts of model violation caused by unincorporated neutral mutational processes. We develop an alternative model that incorporates MNMs and may be helpful in reducing this bias.

evolutionary biology

Speciation as a sieve for ancestral polymorphism

Studying the process of speciation using patterns of genomic divergence between species requires that we understand the determinants of genetic diversity within species. Because sequence diversity in an ancestral population determines the starting point from which divergent populations accumulate differences (Gillespie & Langley 1979), any evolutionary forces that shape diversity within species can have a large impact on measures of divergence between species. These forces include those both decreasing (e.g., selective sweeps (Begun et al. 2007; Cruickshank & Hahn 2014) or background selection (Phung et al. 2016)) and increasing variation (e.g., balancing selection (Charlesworth 2006)). Selection can increase diversity by favoring the maintenance of polymorphism via overdominance, frequency dependence, and heterogeneous selection. Nevertheless, balanced polymorphisms are co ...

evolutionary biology

Why concatenation fails in the anomaly zone

AbstrctGenome-scale sequencing has been of great benefit in recovering species trees, but has not provided final answers. Despite the rapid accumulation of molecular sequences, resolving short and deep branches of the tree of life has remained a challenge, and has prompted the development of new strategies that can make the best use of available data. One such strategy - the concatenation of gene alignments - can be successful when coupled with many tree estimation methods, but has also been shown to fail when there are high levels of incomplete lineage sorting. Here, we focus on the failure of likelihood-based methods in retrieving a rooted, asymmetric four-taxon species tree from concatenated data when the species tree is in or near the anomaly zone - a region of parameter space where the most common gene tree does not match the species tree because of incomplete lineage sorting. First, we use coalescent theory to prove that most informative sites will support the species tree in the anomaly zone, and that as a consequence maximum-parsimony succeeds in recovering the species tree from concatenated data. We further show that maximum-likelihood tree estimation from concatenated data fails both inside and outside the anomaly zone, and that this failure is unconnected to the frequency of the most common gene tree. We provide support for a hypothesis that likelihood-based methods fail in and near the anomaly zone because discordant sites on the species tree have a lower likelihood than those that are discordant on alternative topologies. Our results confirm and extend previous reports of the failure and success of likelihood- and parsimony-based methods, and highlight avenues for future work improving the performance of methods aimed at recovering species tree.

evolutionary biology

The effects of increasing the number of taxa on inferences of molecular convergence

Convergent evolution provides insight into the link between phenotype and genotype. Recently, large-scale comparative studies of convergent evolution have become possible, but researchers are still trying to determine the best way to design these types of analyses. One aspect of molecular convergence studies that has not yet been investigated is how taxonomic sample size affects inferences of molecular convergence. Here we show that increased sample size decreases the amount of inferred molecular convergence associated with the three convergent transitions to a marine environment in mammals. The sampling of more taxa--both with and without the convergent phenotype--reveals that alleles associated only with marine mammals in small datasets are actually more widespread, or are not shared by all marine species. The sampling of more taxa also allows finer resolution of ancestral substitutions, revealing that they are not in fact on lineages leading to solely marine species. We revisit a previous study on marine mammals and find that only 7 of the reported 43 genes with convergent substitutions still show signs of convergence with a larger number of background species. However, 4 of those 7 genes also showed signs of positive selection in the original analysis and may still be good candidates for adaptive convergence. Though our study is framed around the convergence of marine mammals, we expect our conclusions on taxonomic sampling are generalizable to any study of molecular convergence.

genomics

Linking gene expression to unilateral pollen-pistil reproductive barriers

Unilateral incompatibility (UI) is an asymmetric reproductive barrier that unidirectionally prevents gene flow between species and/or populations. UI is characterized by a compatible interaction between partners in one direction, but in the reciprocal cross fertilization fails, generally due to pollen tube rejection by the pistil. Although UI has long been observed in crosses between different species, the underlying molecular mechanisms are only beginning to be characterized. The wild tomato relative Solanum habrochaites provides a unique study system to investigate the molecular basis of this reproductive barrier, as populations within the species exhibit both interspecific and interpopulation UI. Here we used a transcriptomic approach to identify genes in both pollen and pistil tissues that may be probable key players in UI. We confirmed UI at the pollen-pistil level between a self-incompatible population and a self-compatible population of S. habrochaites. A comparison of gene expression between pollinated styles exhibiting the incompatibility response and unpollinated controls revealed only a small number of differentially expressed transcripts. Many more differences in transcript profiles were identified between UI-competent versus UI-compromised reproductive tissues. A number of intriguing candidate genes were highly differentially expressed, including a putative pollen arabinogalactan protein, a stylar Kunitz family protease inhibitor, and a stylar peptide hormone Rapid Alkalinization Factor. Our data also provide transcriptomic evidence that fundamental processes including reactive oxygen species signaling are likely key in UI pollen-pistil interactions between both populations and species. Our transcriptomic analysis highlighted specific genes, including those in ROS signaling pathways that warrant further study in investigations of UI. To our knowledge, this is the first report to identify candidate genes involved in unilateral barriers between populations of the same species.

plant biology