bioRxiv ScienceSearch

Biology subjects

Smith, S. A.

Publications and source records attributed to Smith, S. A..

10 recordsLinked to original sources

Nested phylogenetic conflicts and deep phylogenomics in plants

Studies have demonstrated that pervasive gene tree conflict underlies several important phylogenetic relationships where different species tree methods produce conflicting results. Here, we present a means of dissecting the phylogenetic signal for alternative resolutions within a dataset in order to resolve recalcitrant relationships and, importantly, identify what the dataset is unable to resolve. These procedures extend upon methods for isolating conflict and concordance involving specific candidate relationships and can be used to identify systematic error and disambiguate sources of conflict among species tree inference methods. We demonstrate these on a large phylogenomic plant dataset. Our results support the placement of Amborella as sister to the remaining extant angiosperms, Gnetales as sister to pines, and the monophyly of extant gymnosperms. Several other contentious relationships, including the resolution of relationships within the bryophytes and the eudicots, remain uncertain given the low number of supporting gene trees. To address whether concatenation of filtered genes amplified phylogenetic signal for relationships, we implemented a combinatorial heuristic to test combinability of genes. We found that nested conflicts limited the ability of data filtering methods to fully ameliorate conflicting signal amongst gene trees. These analyses confirmed that the underlying conflicting signal does not support broad concatenation of genes. Our approach provides a means of dissecting a specific dataset to address deep phylogenetic relationships while also identifying the inferential boundaries of the dataset.

evolutionary biology

Evolution of Portulacineae marked by gene tree conflict and gene family expansion associated with adaptation to harsh environments

Several plant lineages have evolved adaptations that allow survival in extreme and harsh environments including many within the plant clade Portulacineae (Caryophyllales) such as the Cactaceae, Didiereaceae of Madagascar, and high altitude Montiaceae. Here, using newly generated transcriptomic data, we reconstructed the phylogeny of Portulacineae and examine potential correlates between molecular evolution within this clade and adaptation to harsh environments. Our phylogenetic results were largely congruent with previous analyses, but we identified several early diverging nodes characterized by extensive gene tree conflict. For particularly contentious nodes, we presented detailed information about the phylogenetic signal for alternative relationships. We also analyzed the frequency of gene duplications, confirmed previously identified whole genome duplications (WGD), and identified a previously unidentified WGD event within the Didiereaceae. We found that the WGD events were typically associated with shifts in climatic niche and did not find a direct association with WGDs and diversification rate shifts. Diversification shifts occurred within the Portulacaceae, Cactaceae, and Anacampserotaceae, and while these did not experience WGDs, the Cactaceae experienced extensive gene duplications. We examined gene family expansion and molecular evolutionary patterns with a focus on genes associated with environmental stress responses and found evidence for significant gene family expansion in genes with stress adaptation and clades found in extreme environments. These results provide important directions for further and deeper examination of the potential links between molecular evolutionary patterns and adaptation to harsh environments.

evolutionary biology

Genome wide association analysis identifies genetic variants associated with reproductive variation across domestic dog breeds and uncovers links to domestication

The diversity of eutherian reproductive strategies has led to variation in many traits, such as number of offspring, age of reproductive maturity, and gestation length. While reproductive trait variation has been extensively investigated and is well established in mammals, the genetic loci contributing to this variation remain largely unknown. The domestic dog, Canis lupus familiaris is a powerful model for studies of the genetics of inherited disease due to its unique history of domestication. To gain insight into the genetic basis of reproductive traits across domestic dog breeds, we collected phenotypic data for four traits - cesarean section rate (n = 97 breeds), litter size (n = 60), stillbirth rate (n = 57), and gestation length (n = 23) - from primary literature and breeders handbooks. By matching our phenotypic data to genomic data from the Cornell Veterinary Biobank, we performed genome wide association analyses for these four reproductive traits, using body mass and kinship among breeds as co-variates. We identified 14 genome-wide significant associations between these traits and genetic loci, including variants near CACNA2D3 with gestation length, MSRB3 with litter size, SMOC2 with cesarean section rate, MITF with litter size and still birth rate, KRT71 with cesarean section rate, litter size, and stillbirth rate, and HTR2C with stillbirth rate. Some of these loci, such as CACNA2D3 and MSRB3, have been previously implicated in human reproductive pathologies. Many of the variants that we identified have been previously associated with domestication-related traits, including brachycephaly (SMOC2), coat color (MITF), coat curl (KRT71), and tameness (HTR2C). These results raise the hypothesis that the artificial selection that gave rise to dog breeds also shaped the observed variation in their reproductive traits. Overall, our work establishes the domestic dog as a system for studying the genetics of reproductive biology and disease.

evolutionary biology

Pan-arthropod analysis reveals somatic piRNAs as an ancestral TE defence

In animals, small RNA molecules termed PIWI-interacting RNAs (piRNAs) silence transposable elements (TEs), protecting the germline from genomic instability and mutation. piRNAs have been detected in the soma in a few animals, but these are believed to be specific adaptations of individual species. Here, we report that somatic piRNAs were likely present in the ancestral arthropod more than 500 million years ago. Analysis of 20 species across the arthropod phylum suggests that somatic piRNAs targeting TEs and mRNAs are common among arthropods. The presence of an RNA-dependent RNA polymerase in chelicerates (horseshoe crabs, spiders, scorpions) suggests that arthropods originally used a plant-like RNA interference mechanism to silence TEs. Our results call into question the view that the ancestral role of the piRNA pathway was to protect the germline and demonstrate that small RNA silencing pathways have been repurposed for both somatic and germline functions throughout arthropod evolution.

evolutionary biology

Quartet Sampling distinguishes lack of support from conflicting support \newline in the plant tree of life

Premise of the StudyPhylogenetic support has been difficult to evaluate within the plant tree of life partly due to the difficulty of distinguishing conflicted versus poorly informed branches. As datasets continue to expand in both breadth and depth, new support measures are needed that are more efficient and informative.\n\nMethodsWe describe the Quartet Sampling (QS) method, a quartet-based evaluation system that synthesizes several phylogenetic and genomic analytical approaches. QS characterizes discordance in large-sparse and genome-wide datasets, overcoming issues of alignment sparsity and distinguishing strong conflict from weak support. We test QS with simulations and recent plant phylogenies inferred from variously sized datasets.\n\nKey ResultsQS scores demonstrate convergence with increasing replicates and are not strongly affected by branch depth. Patterns of QS support from different phylogenies leads to a coherent understanding of ancestral branches defining key disagreements, including the relationships of Ginkgo to cycads, magnoliids to monocots and eudicots, and mosses to liverworts. The relationships of ANA grade angiosperms, major monocot groups, bryophytes, and fern families are likely highly discordant in their evolutionary histories, rather than poorly informed. QS can also detect discordance due to introgression in phylogenomic data.\n\nConclusionsThe QS method represents an efficient and effective synthesis of phylogenetic tests that offer more comprehensive and specific information on branch support than conventional measures. The QS method corroborates growing evidence that phylogenomic investigations that incorporate discordance testing are warranted to reconstruct the complex evolutionary histories surrounding in particular ANA grade angiosperms, monocots, and non-vascular plants.

evolutionary biology

Widespread paleopolyploidy, gene tree conflict, and recalcitrant relationships among the carnivorous Caryophyllales

O_LIThe carnivorous members of the large, hyperdiverse Caryophyllales (e.g. Venus flytrap, sundews and Nepenthes pitcher plants) represent perhaps the oldest and most diverse lineage of carnivorous plants. However, despite numerous studies seeking to elucidate their evolutionary relationships, the early-diverging relationships remain unresolved.\nC_LIO_LITo explore the utility of phylogenomic data sets for resolving relationships among the carnivorous Caryophyllales, we sequenced ten transcriptomes, including all the carnivorous genera except those in the rare West African liana family (Dioncophyllaceae). We used a variety of methods to infer the species tree, examine gene tree conflict and infer paleopolyploidy events.\nC_LIO_LIPhylogenomic analyses support the monophyly of the carnivorous Caryophyllales, with an origin of 68-83 mya. In contrast to previous analyses recover the remaining non-core Caryophyllales as non-monophyletic, although there are multiple reasons this result may be spurious and node supporting this relationship contains a significant amount gene tree discordance. We present evidence that the clade contains at least seven independent paleopolyploidy events, previously debated nodes from the literature have high levels of gene tree conflict, and taxon sampling influences topology even in a phylogenomic data set.\nC_LIO_LIOur data demonstrate the importance of carefully considering gene tree conflict and taxon sampling in phylogenomic analyses. Moreover, they provide a remarkable example of the propensity for paleopolyploidy in angiosperms, with at least seven such events in a clade of less than 2500 species.\nC_LI

evolutionary biology

Site and gene-wise likelihoods unmask influential outliers in phylogenomic analyses

Recent studies have demonstrated that conflict is common among gene trees in phylogenomic studies, and that less than one percent of genes may ultimately drive species tree inference in supermatrix analyses. Here, we examined two datasets where supermatrix and coalescent-based species trees conflict. We identified two highly influential \"outlier\" genes in each dataset. When removed from each dataset, the inferred supermatrix trees matched the topologies obtained from coalescent analyses. We also demonstrate that, while the outlier genes in the vertebrate dataset have been shown in a previous study to be the result of errors in orthology detection, the outlier genes from a plant dataset did not exhibit any obvious systematic error and therefore may be the result of some biological process yet to be determined. While topological comparisons among a small set of alternate topologies can be helpful in discovering outlier genes, they can be limited in several ways, such as assuming all genes share the same topology. Coalescent species tree methods relax this assumption but do not explicitly facilitate the examination of specific edges. Coalescent methods often also assume that conflict is the result of incomplete lineage sorting (ILS). Here we explored a framework that allows for quickly examining alternative edges and support for large phylogenomic datasets that does not assume a single topology for all genes. For both datasets, these analyses provided detailed results confirming the support for coalescent-based topologies. This framework suggests that we can improve our understanding of the underlying signal in phylogenomic datasets by asking more targeted edge-based questions.

evolutionary biology

Missing the point (estimate): Bayesian and likelihood phylogenetic reconstructions of morphological characters produce generally concordant inferences. A comment on Puttick et al.

Puttick et al. [1] performed a simulation study to compare accuracy among methods of inferring phylogeny from discrete morphological characters. They report that a Bayesian implementation of the Mk model [2] was most accurate (but with low resolution), while a maximum likelihood (ML) implementation of the same model was least accurate. They conclude by strongly advocating that Bayesian implementations of the Mk model should be the default method of analysis for such data. While we appreciate the authors attempt to investigate the accuracy of alternative methods of analysis, their conclusion is based on an inappropriate comparison of the ML point estimate, which does not consider confidence, with the Bayesian consensus, which incorporates estimation credibility into the summary tree. Using simulation, we demonstrate that ML and Bayesian estimates are concordant when confidence and credibility are comparably reflected in summary trees, a result expected from statistical theory. We therefore disagree with the conclusions of PEA and consider their prescription of any default method to be poorly founded. Instead, we recommend caution and thoughtful consideration of the model or method being applied to a morphological dataset.

evolutionary biology

The Past Sure Is Tense: On Interpreting Phylogenetic Divergence Time Estimates

Divergence time estimation -- the calibration of a phylogeny to geological time -- is a integral first step in modelling the tempo of biological evolution (traits and lineages). However, despite increasingly sophisticated methods to infer divergence times from molecular genetic sequences, the estimated age of many nodes across the tree of life contrast significantly and consistently with timeframes conveyed by the fossil record. This is perhaps best exemplified by crown angiosperms, where molecular clock (Triassic) estimates predate the oldest (Early Cretaceous) undisputed angiosperm fossils by tens of millions of years or more. While the incompleteness of the fossil record is a common concern, issues of data limitation and model inadequacy are viable (if underexplored) alternative explanations. In this vein, Beaulieu et al. (2015) convincingly demonstrated how methods of divergence time inference can be misled by both (i) extreme state-dependent molecular substitution rate heterogeneity and (ii) biased sampling of representative major lineages. While these (essentially model-violation) results are robust (and probably common in empirical data sets), we note a further alternative: that the configuration of the statistical inference problem alone (i.e., the parameters, their relationships, and associated priors) precluded the reconstruction of the paleontological timeframe for the crown age of angiosperms. We demonstrate, through sampling from the joint prior (formed by combining the tree (diversification) prior with the various calibration densities specified for fossil-calibrated nodes), that with no data present at all, an Early Cretaceous crown angiosperms is rejected (i.e., has essentially zero probability). More worrisome, however, is that for the 24 nodes calibrated by fossils, almost all have indistinguishable marginal prior and posterior age distributions, indicating an absence of relevant information in the data. Given that these calibrated nodes are strategically placed in disparate regions of the tree, they essentially anchor the tree scaffold, and so the posterior inference for the tree as a whole is largely determined by the pseudo-data present in the (often arbitrary) calibration densities. We recommend, as for any Bayesian analysis, that marginal prior and posterior distributions be carefully compared, especially for parameters of direct interest. Finally, we note that the results presented here do not refute the biological modelling concerns identified by Beaulieu et al. (2015). Both sets of issues remain apposite to the goals of accurate divergence time estimation, and only by considering them in tandem can we move forward more confidently. [marginal priors; information content; diptych; divergence time estimation; fossil record; BEAST; angiosperms.]

evolutionary biology

Inhibition of microbial biofuel production in drought stressed switchgrass hydrolysate

BackgroundInterannual variability in precipitation, particularly drought, can affect lignocellulosic crop biomass yields and composition, and is expected to increase biofuel yield variability. However, the effect of precipitation on downstream fermentation processes has never been directly characterized. In order to investigate the impact of interannual climate variability on biofuel production, corn stover and switchgrass were collected during three years with significantly different precipitation profiles, representing a major drought year (2012) and two years with average precipitation for the entire season (2010 and 2013). All feedstocks were AFEX (ammonia fiber expansion)-pretreated, enzymatically hydrolyzed, and the hydrolysates separately fermented using xylose-utilizing strains of Saccharomyces cerevisiae and Zymomonas mobilis. A chemical genomics approach was also used to evaluate the growth of yeast mutants in the hydrolysates.\n\nResultsWhile most corn stover and switchgrass hydrolysates were readily fermented, growth of S. cerevisiae was completely inhibited in hydrolysate generated from drought stressed switchgrass. Based on chemical genomics analysis, yeast strains deficient in genes related to protein trafficking within the cell were significantly more resistant to the drought year switchgrass hydrolysate. Detailed biomass and hydrolysate characterization revealed that switchgrass accumulated greater concentrations of soluble sugars in response to the drought and these sugars were subsequently degraded to pyrazines and imidazoles during ammonia-based pretreatment. When added ex situ to normal switchgrass hydrolysate, imidazoles and pyrazines caused anaerobic growth inhibition of S. cerevisiae.\n\nConclusionsIn response to the osmotic pressures experienced during drought stress, plants accumulate soluble sugars that are susceptible to degradation during chemical pretreatments. For ammonia-based pretreatment these sugars degrade to imidazoles and pyrazines. These compounds contribute to S. cerevisiae growth inhibition in drought year switchgrass hydrolysate. This work discovered that variation in environmental conditions during the growth of bioenergy crops could have significant detrimental effects on fermentation organisms during biofuel production. These findings are relevant to regions where climate change is predicted to cause an increased incidence of drought and to marginal lands with poor water holding capacity, where fluctuations in soil moisture may trigger frequent drought stress response in lignocellulosic feedstocks.

microbiology