bioRxiv ScienceSearch

Biology subjects

Brown, J. W.

Publications and source records attributed to Brown, J. W..

9 recordsLinked to original sources

The choice of tree prior and molecular clock does not substantially affect phylogenetic inferences of diversification rates

Comparative methods allow researchers to make inferences about evolutionary processes and patterns from phylogenetic trees. In Bayesian phylogenetics, estimating a phylogeny requires specifying priors on parameters characterizing the branching process and rates of substitution among lineages, in addition to others. However, the effect that the selection of these priors has on the inference of comparative parameters has not been thoroughly investigated. Such uncertainty may systematically bias phylogenetic reconstruction and, subsequently, parameter estimation. Here, we focus on the impact of priors in Bayesian phylogenetic inference and evaluate how they affect the estimation of parameters in macroevolutionary models of lineage diversification. Specifically, we use BEAST to simulate trees under combinations of tree priors and molecular clocks, simulate sequence data, estimate trees, and estimate diversification parameters (e.g., speciation rates and extinction rates) from these trees. When substitution rate heterogeneity is large, parameter estimates deviate substantially from those estimated under the simulation conditions when not captured by an appropriate choice of relaxed molecular clock. However, in general, we find that the choice of tree prior and molecular clock has relatively little impact on the estimation of diversification rates insofar as the sequence data are sufficiently informative and substitution rate heterogeneity among lineages is low-to-moderate.

evolutionary biology

Evolution of Portulacineae marked by gene tree conflict and gene family expansion associated with adaptation to harsh environments

Several plant lineages have evolved adaptations that allow survival in extreme and harsh environments including many within the plant clade Portulacineae (Caryophyllales) such as the Cactaceae, Didiereaceae of Madagascar, and high altitude Montiaceae. Here, using newly generated transcriptomic data, we reconstructed the phylogeny of Portulacineae and examine potential correlates between molecular evolution within this clade and adaptation to harsh environments. Our phylogenetic results were largely congruent with previous analyses, but we identified several early diverging nodes characterized by extensive gene tree conflict. For particularly contentious nodes, we presented detailed information about the phylogenetic signal for alternative relationships. We also analyzed the frequency of gene duplications, confirmed previously identified whole genome duplications (WGD), and identified a previously unidentified WGD event within the Didiereaceae. We found that the WGD events were typically associated with shifts in climatic niche and did not find a direct association with WGDs and diversification rate shifts. Diversification shifts occurred within the Portulacaceae, Cactaceae, and Anacampserotaceae, and while these did not experience WGDs, the Cactaceae experienced extensive gene duplications. We examined gene family expansion and molecular evolutionary patterns with a focus on genes associated with environmental stress responses and found evidence for significant gene family expansion in genes with stress adaptation and clades found in extreme environments. These results provide important directions for further and deeper examination of the potential links between molecular evolutionary patterns and adaptation to harsh environments.

evolutionary biology

What drives results in Bayesian morphological clock analyses?

Recently, approaches that estimate species divergence times using fossil taxa and models of morphological evolution have exploded in popularity. These methods incorporate diverse biological and geological information to inform posterior reconstructions, and have been applied to several high-profile clades to positive effect. However, there are important examples where morphological data are misleading, resulting in unrealistic age estimates. While several studies have demonstrated that these approaches can be robust and internally consistent, the causes and limitations of these patterns remain unclear. In this study, we dissect signal in Bayesian dating analyses of three mammalian clades. For two of the three examples, we find that morphological characters provide little information regarding divergence times as compared to geological range information, with posterior estimates largely recapitulating those recovered under the prior. However, in the cetacean dataset, we find that morphological data do appreciably inform posterior divergence time estimates. We supplement these empirical analyses with a set of simulations designed to explore the efficiency and limitations of binary and 3-state character data in reconstructing node ages. Our results demonstrate areas of both strength and weakness for morphological clock analyses, and help to outline conditions under which they perform best and, conversely, when they should be eschewed in favour of purely geological approaches.

paleontology

Quartet Sampling distinguishes lack of support from conflicting support \newline in the plant tree of life

Premise of the StudyPhylogenetic support has been difficult to evaluate within the plant tree of life partly due to the difficulty of distinguishing conflicted versus poorly informed branches. As datasets continue to expand in both breadth and depth, new support measures are needed that are more efficient and informative.\n\nMethodsWe describe the Quartet Sampling (QS) method, a quartet-based evaluation system that synthesizes several phylogenetic and genomic analytical approaches. QS characterizes discordance in large-sparse and genome-wide datasets, overcoming issues of alignment sparsity and distinguishing strong conflict from weak support. We test QS with simulations and recent plant phylogenies inferred from variously sized datasets.\n\nKey ResultsQS scores demonstrate convergence with increasing replicates and are not strongly affected by branch depth. Patterns of QS support from different phylogenies leads to a coherent understanding of ancestral branches defining key disagreements, including the relationships of Ginkgo to cycads, magnoliids to monocots and eudicots, and mosses to liverworts. The relationships of ANA grade angiosperms, major monocot groups, bryophytes, and fern families are likely highly discordant in their evolutionary histories, rather than poorly informed. QS can also detect discordance due to introgression in phylogenomic data.\n\nConclusionsThe QS method represents an efficient and effective synthesis of phylogenetic tests that offer more comprehensive and specific information on branch support than conventional measures. The QS method corroborates growing evidence that phylogenomic investigations that incorporate discordance testing are warranted to reconstruct the complex evolutionary histories surrounding in particular ANA grade angiosperms, monocots, and non-vascular plants.

evolutionary biology

Disparity, Diversity, And Duplications In The Caryophyllales

O_LIThe role whole genome duplication (WGD) plays in the history of lineages is actively debated. WGDs have been associated with advantages including superior colonization, various adaptations, and increased effective population size. However, the lack of a comprehensive mapping of WGDs within a major plant clade has led to uncertainty regarding the potential association of WGDs and higher diversification rates.\nC_LIO_LIUsing seven chloroplast and nuclear ribosomal genes, we constructed a phylogeny of 5,036 species of Caryophyllales, representing nearly half of the extant species. We phylogenetically mapped putative WGDs as identified from analyses on transcriptomic and genomic data and analyzed these in conjunction with shifts in climatic niche and lineage diversification rate.\nC_LIO_LIThirteen putative WGDs and twenty-seven diversification shifts could be mapped onto the phylogeny. Of these, four WGDs were concurrent with diversification shifts, with other diversification shifts occurring at more recent nodes than WGDs. Five WGDs were associated with shifts to colder climatic niches.\nC_LIO_LIWhile we find that many diversification shifts occur after WGDs it is difficult to consider diversification and duplication to be tightly correlated. Our findings suggest that duplications may often occur along with shifts in either diversification rate, climatic niche, or rate of evolution.\nC_LI

evolutionary biology

Site and gene-wise likelihoods unmask influential outliers in phylogenomic analyses

Recent studies have demonstrated that conflict is common among gene trees in phylogenomic studies, and that less than one percent of genes may ultimately drive species tree inference in supermatrix analyses. Here, we examined two datasets where supermatrix and coalescent-based species trees conflict. We identified two highly influential \"outlier\" genes in each dataset. When removed from each dataset, the inferred supermatrix trees matched the topologies obtained from coalescent analyses. We also demonstrate that, while the outlier genes in the vertebrate dataset have been shown in a previous study to be the result of errors in orthology detection, the outlier genes from a plant dataset did not exhibit any obvious systematic error and therefore may be the result of some biological process yet to be determined. While topological comparisons among a small set of alternate topologies can be helpful in discovering outlier genes, they can be limited in several ways, such as assuming all genes share the same topology. Coalescent species tree methods relax this assumption but do not explicitly facilitate the examination of specific edges. Coalescent methods often also assume that conflict is the result of incomplete lineage sorting (ILS). Here we explored a framework that allows for quickly examining alternative edges and support for large phylogenomic datasets that does not assume a single topology for all genes. For both datasets, these analyses provided detailed results confirming the support for coalescent-based topologies. This framework suggests that we can improve our understanding of the underlying signal in phylogenomic datasets by asking more targeted edge-based questions.

evolutionary biology

So many genes, so little time: comments on divergence-time estimation in the genomic era

Phylogenomic datasets have been successfully used to address questions involving evolutionary relationships, patterns of genome structure, signatures of selection, and gene and genome duplications. However, despite the recent explosion in genomic and transcriptomic data, the utility of these data sources for efficient divergence-time inference remains unexamined. Phylogenomic datasets pose two distinct problems for divergence-time estimation: (i) the volume of data makes inference of the entire dataset intractable, and (ii) the extent of underlying topological and rate heterogeneity across genes makes model mis-specification a real concern. \"Gene shopping\", wherein a phylogenomic dataset is winnowed to a set of genes with desirable properties, represents an alternative approach that holds promise in alleviating these issues. We implemented an approach for phylogenomic datasets (available in SortaDate) that filters genes by three criteria: (i) clock-likeness, (ii) reasonable tree length (i.e., discernible information content), and (iii) least topological conflict with a focal species tree (presumed to have already been inferred). Such a winnowing procedure ensures that errors associated with model (both clock and topology) mis-specification are minimized, therefore reducing error in divergence-time estimation. We demonstrated the efficacy of this approach through simulation and applied it to published animal (Aves, Diplopoda, and Hymenoptera) and plant (carnivorous Caryophyllales, broad Caryophyllales, and Vitales) phylogenomic datasets. By quantifying rate heterogeneity across both genes and lineages we found that every empirical dataset examined included genes with clock-like, or nearly clock-like, behavior. Moreover, many datasets had genes that were clock-like, exhibited reasonable evolutionary rates, and were mostly compatible with the species tree. We identified overlap in age estimates when analyzing these filtered genes under strict clock and uncorrelated lognormal (UCLN) models. However, this overlap was often due to imprecise estimates from the UCLN model. We find that \"gene shopping\" can be an efficient approach to divergence-time inference for phylogenomic datasets that may otherwise be characterized by extensive gene tree heterogeneity.

evolutionary biology

Missing the point (estimate): Bayesian and likelihood phylogenetic reconstructions of morphological characters produce generally concordant inferences. A comment on Puttick et al.

Puttick et al. [1] performed a simulation study to compare accuracy among methods of inferring phylogeny from discrete morphological characters. They report that a Bayesian implementation of the Mk model [2] was most accurate (but with low resolution), while a maximum likelihood (ML) implementation of the same model was least accurate. They conclude by strongly advocating that Bayesian implementations of the Mk model should be the default method of analysis for such data. While we appreciate the authors attempt to investigate the accuracy of alternative methods of analysis, their conclusion is based on an inappropriate comparison of the ML point estimate, which does not consider confidence, with the Bayesian consensus, which incorporates estimation credibility into the summary tree. Using simulation, we demonstrate that ML and Bayesian estimates are concordant when confidence and credibility are comparably reflected in summary trees, a result expected from statistical theory. We therefore disagree with the conclusions of PEA and consider their prescription of any default method to be poorly founded. Instead, we recommend caution and thoughtful consideration of the model or method being applied to a morphological dataset.

evolutionary biology

The Past Sure Is Tense: On Interpreting Phylogenetic Divergence Time Estimates

Divergence time estimation -- the calibration of a phylogeny to geological time -- is a integral first step in modelling the tempo of biological evolution (traits and lineages). However, despite increasingly sophisticated methods to infer divergence times from molecular genetic sequences, the estimated age of many nodes across the tree of life contrast significantly and consistently with timeframes conveyed by the fossil record. This is perhaps best exemplified by crown angiosperms, where molecular clock (Triassic) estimates predate the oldest (Early Cretaceous) undisputed angiosperm fossils by tens of millions of years or more. While the incompleteness of the fossil record is a common concern, issues of data limitation and model inadequacy are viable (if underexplored) alternative explanations. In this vein, Beaulieu et al. (2015) convincingly demonstrated how methods of divergence time inference can be misled by both (i) extreme state-dependent molecular substitution rate heterogeneity and (ii) biased sampling of representative major lineages. While these (essentially model-violation) results are robust (and probably common in empirical data sets), we note a further alternative: that the configuration of the statistical inference problem alone (i.e., the parameters, their relationships, and associated priors) precluded the reconstruction of the paleontological timeframe for the crown age of angiosperms. We demonstrate, through sampling from the joint prior (formed by combining the tree (diversification) prior with the various calibration densities specified for fossil-calibrated nodes), that with no data present at all, an Early Cretaceous crown angiosperms is rejected (i.e., has essentially zero probability). More worrisome, however, is that for the 24 nodes calibrated by fossils, almost all have indistinguishable marginal prior and posterior age distributions, indicating an absence of relevant information in the data. Given that these calibrated nodes are strategically placed in disparate regions of the tree, they essentially anchor the tree scaffold, and so the posterior inference for the tree as a whole is largely determined by the pseudo-data present in the (often arbitrary) calibration densities. We recommend, as for any Bayesian analysis, that marginal prior and posterior distributions be carefully compared, especially for parameters of direct interest. Finally, we note that the results presented here do not refute the biological modelling concerns identified by Beaulieu et al. (2015). Both sets of issues remain apposite to the goals of accurate divergence time estimation, and only by considering them in tandem can we move forward more confidently. [marginal priors; information content; diptych; divergence time estimation; fossil record; BEAST; angiosperms.]

evolutionary biology