bioRxiv ScienceSearch

Biology subjects

Laura Kubatko

Publications and source records attributed to Laura Kubatko.

6 recordsLinked to original sources

Genetic Diversity, Population Structure and Species Delimitation of Trialeurodes vaporariorum (Greenhouse whitefly)

Genetic diversity within Trialeurodes vaporariorum (Westwood, 1856) remains largely unexplored, particularly within regions of Sub-Saharan Africa. In this study, T. vaporariorum samples were obtained from three locations in Kenya: Katumani, Kiambu and Kajiado counties. DNA extraction, PCR and Sanger sequencing were carried out on ~750 bp fragment of the mitochondria cytochrome c oxidase I (COI) gene from individual whiteflies. In addition, global populations were assessed and 19 haplotypes were identified, with three main haplotypes (Hp_19, Hp_10, Hp_011) circulating within Kenya. Measures of genetic diversity among T. vaporariorum populations resulted in haplotype diversity of 0.411, nucleotide diversity 0.00096, and Tajimas D -0. 30315, (P>0.10). Analysis of population structure across global sequences using Structurama indicated one population globally, with posterior probability of 0.72. Bayesian and maximum likelihood phylogenetic analysis gave support for two clades (Clade I = an admixed global population and Clade II = subset of Kenyan and 1 Greek sequence). Species delimitation between the two clades was assessed by four parameters; posterior probability, Kimuras two parameter (K2P), Rodrigos P (Randomly distinct) and Rosenbergs reciprocal monophyly (P(AB). The two clades within the phylogenetic tree showed evidence of distinctness based on; Kimura two parameters (K2P) (p = -1.21E-01), Rodrigos P (RD) (p =0.05) and Rosenbergs P(AB) (p = 2.3E -13). Overall, low genetic diversity within the Kenyan samples is a likely indicator of recent population expansion and colonization with this region and plausible signs of species complex formation in Sub-Saharan Africa.

Evolutionary Biology

Characterization by Next Generation Sequencing Reveals the Molecular Mechanisms Driving the Faster Evolutionary rate of Cassava brown streak virus Compared with Ugandan cassava brown streak virus

Cassava is a major staple food for about 800 million people in the tropics and subGtropical regions of the world. Production of cassava is significantly hampered by cassava brown streak disease (CBSD), which is caused by Cassava brown streak virus (CBSV) and Ugandan cassava brown streak virus (UCBSV). The disease is suppressing cassava yields in eastern Africa at an alarming rate. Previous studies have documented that CBSV is more devastating than UCBSV because it more readily infects both susceptible and tolerant cassava cultivars, resulting in greater yield losses. Using whole genome sequences from NGS data, we produced the first coalescentGbased species tree estimate for CBSV and UCBSV. This species framework led to the finding that CBSV has a faster rate of evolution when compared with UCBSV. Furthermore, we have discovered that in CBSV, nonsynonymous substitutions are more predominant than synonymous substitution and occur across the entire genome. All comparative analyses between CBSV and UCBSV presented here suggest that CBSV may be outsmarting the cassava immune system, thus making it more devastating and harder to control.

Evolutionary Biology

An Invariants-based Method for Efficient Identification of Hybrid Species From Large-scale Genomic Data

Coalescent-based species tree inference has become widely used in the analysis of genome-scale multilocus and SNP datasets when the goal is inference of a species-level phylogeny. However, numerous evolutionary processes are known to violate the assumptions of a coalescence-only model and complicate inference of the species tree. One such process is hybrid speciation, in which a species shares its ancestry with two distinct species. Although many methods have been proposed to detect hybrid speciation, only a few have considered both hybridization and coalescence in a unified framework, and these are generally limited to the setting in which putative hybrid species must be identified in advance. Here we propose a method that can examine genome-scale data for a large number of taxa and detect those taxa that may have arisen via hybridization, as well as their potential \"parental\" taxa. The method is based on a model that considers both coalescence and hybridization together, and uses phylogenetic invariants to construct a test that scales well in terms of computational time for both the number of taxa and the amount of sequence data. We test the method using simulated data for up 20 taxa and 100,000bp, and find that the method accurately identifies both recent and ancient hybrid species in less than 30 seconds. We apply the method to two empirical datasets, one composed of Sistrurus rattlesnakes for which hybrid speciation is not supported by previous work, and one consisting of several species of Heliconius butterflies for which some evidence of hybrid speciation has been previously found.

Evolutionary Biology

Distribution of gene tree histories under the coalescent model with gene flow

We propose a coalescent model for three species that allows gene flow between both pairs of sister populations. The model is designed to analyze multilocus genomic sequence alignments, with one sequence sampled from each of the three species. The model is formulated using a Markov chain representation, which allows use of matrix exponentiation to compute analytical expressions for the probability density of gene tree genealogies. The gene tree history distribution as well as the gene tree topology distribution under this coalescent model with gene flow are then calculated via numerical integration. We analyze the model to compare the distributions of gene tree topologies and gene tree histories for species trees with differing effective population sizes and gene flow rates. Our results suggest conditions under which the species tree and associated parameters are not identifiable from the gene tree topology distribution when gene flow is present, but indicate that the gene tree history distribution may identify the species tree and associated parameters. Thus, the gene tree history distribution can be used to infer parameters such as the ancestral effective population sizes and the rates of gene flow in a maximum likelihood (ML) framework. We conduct computer simulations to evaluate the performance of our method in estimating these parameters, and we apply our method to an Afrotropical mosquito data set (Fontaine et al., 2015) to demonstrate the usefulness of our method for the analysis of empirical data.

Evolutionary Biology

A Distance Method to Reconstruct Species Trees In the Presence of Gene Flow

One of the central tasks in evolutionary biology is to reconstruct the evolutionary relationships among species from sequence data, particularly from multilocus data. In the last ten years, many methods have been proposed to use the variance in the gene histories to estimate species trees by explicitly modeling deep coalescence. However, gene flow, another process that may produce gene history variance, has been less studied. In this paper, we propose a simple yet innovative method for species trees estimation in the presence of gene flow. Our method, called STEST (Species Tree Estimation from Speciation Times), constructs species tree estimates from pairwise speciation time or species divergence time estimates. By using methods that estimate speciation times in the presence of gene flow, (for example, M1 (Yang 2010) or SIM3s (Zhu and Yang 2012)), STEST is able to estimate species trees from data subject to gene flow. We develop two methods, called STEST (M1) and STEST (SIM3s), for this purpose. Additionally, we consider the method STEST (M0), which instead uses the M0 method (Yang 2002), a coalescent-based method that does not assume gene flow, to estimate speciation times. It is therefore devised to estimate species trees in the absence of gene flow. Our simulation studies show that STEST (M0) outperforms STEST(M1), STEST (SIM3s) and STEM in terms of estimation accuracy and outperfroms *BEAST in terms of running time when the degree of gene flow is small. STEST (M1) outperforms STEST (M0), STEST (SIM3s), STEM and *BEAST in term of estimation accuracy when the degree of gene flow is large. An empirical data set analyzed by these methods gives species tree estimates that are consistent with the previous results.

Evolutionary Biology

A codon model of nucleotide substitution with selection on synonymous codon usage

The quality of phylogenetic inference made from protein-coding genes depends, in part, on the realism with which the codon substitution process is modeled. Here we propose a new mechanistic model that combines the standard M0 substitution model of Yang (1997) with a simplified model from Gilchrist (2007) that includes selection on synonymous substitutions as a function of codon-specific nonsense error rates. We tested the newly proposed model by applying it to 104 protein-coding genes in brewer's yeast, and compared the fit of the new model to the standard M0 model and to the mutation-selection model of Yang and Nielsen (2008) using the AIC. Our new model provided significantly better fit in approximately 85% of the cases considered for the basic M0 model and in approximately 25% of the cases for the M0 model with estimated codon frequencies, but only in a few cases when the mutation-selection model was considered. However, our model includes a parameter that can be interpreted as a measure of the rate of protein production, and the estimates of this parameter were highly correlated with an independent measure of protein production for the yeast genes considered here. Finally, we found that in some cases the new model led to the preference of a different phylogeny for a subset of the genes considered, indicating that substitution model choice may have an impact on the estimated phylogeny.

Evolutionary Biology