bioRxiv Science⌕ Search

Biology subjects

Chan, Y.-b.

Publications and source records attributed to Chan, Y.-b..

4 recordsLinked to original sources

The accuracy of species tree inference under gene tree dependence

When inferring the evolutionary history of species and the genes they contain, the phylogenetic trees of genes can be different from those of the species and to each other, due to a variety of causes, including incomplete lineage sorting. We often wish to infer the species tree, but only reconstruct the gene trees from sequences. We then combine the gene trees to produce a species tree; methods to do this are known as summary methods, of which ASTRAL is currently among the most popular. ASTRAL has been shown to be accurate in many practical scenarios through extensive simulations. However, these simulations generally assume that the input gene trees are independent of each other (infinite recombination between loci). This is known to be unrealistic, as genes that are close to each other on the chromosome (or are co-evolving) have dependent phylogenies. In this paper, we develop a model for generating dependent gene trees within a species tree, based on the coalescent with recombination. We then use these trees as input to ASTRAL to reassess its accuracy for dependent gene trees. Our results allow us to evaluate the impact of any level of dependence on the accuracy of ASTRAL, both when gene trees are known and estimated from sequences. We find that a fixed amount of dependence reduces the effective sample size by a constant factor. In current phylogenomic datasets, loci are generally sampled at large genomic distances to reduce gene tree dependence, thereby limiting the number of genes available for inference. However, full independence between genes is not required for accurate species tree estimation, and excluding gene trees may reduce inference accuracy. This creates a trade-off between the number of genes used and the degree of gene tree dependence. We therefore propose a method to identify the minimum genomic separation required to maintain satisfactory inference accuracy.

evolutionary biology↗

Estimating evolutionary and demographic parameters via ARG-derived IBD

Inference of demographic and evolutionary parameters from a sample of genome sequences often proceeds by first inferring identical-by-descent (IBD) genome segments. By exploiting efficient data encoding based on the ancestral recombination graph (ARG), we obtain three major advantages over current approaches: (i) no need to impose a length threshold on IBD segments, (ii) IBD can be defined without the hard-to-verify requirement of no recombination, and (iii) computation time can be reduced with little loss of statistical efficiency using only the IBD segments from a set of sequence pairs that scales linearly with sample size. We first demonstrate powerful inferences when true IBD information is available from simulated data. For IBD inferred from real data, we propose an approximate Bayesian computation inference algorithm and use it to show that poorly-inferred short IBD segments can improve estimation precision. We show estimation precision similar to a previously-published estimator despite a 4 000-fold reduction in data used for inference. Computational cost limits model complexity in our approach, but we are able to incorporate unknown nuisance parameters and model misspecification, still finding improved parameter inference. Author summarySamples of genome sequences can be informative about the history of the population from which they were drawn, and about mutation and other processes that led to the observed sequences. However, obtaining reliable inferences is challenging, because of the complexity of the underlying processes and the large amounts of sequence data that are often now available. A common approach to simplifying the data is to use only genome segments that are very similar between two sequences, called identical-by-descent (IBD). The longer the IBD segment the more informative about recent shared ancestry, and current approaches restrict attention to IBD segments above a length threshold. We instead are able to use IBD segments of any length, allowing us to extract much more information from the sequence data. To reduce the computation burden we identify subsets of the available sequence pairs that lead to little information loss. Our approach exploits recent advances in inferring aspects of the ancestral recombination graph (ARG) underlying the sample of sequences. Computational cost still limits the size and complexity of problems our method can handle, but where feasible we obtain dramatic improvements in the power of inferences.

genetics↗

A paradoxical population structure of var DBLα types in Africa

The var multigene family encodes the P. falciparum erythrocyte membrane protein 1 (PfEMP1), which is important in host-parasite interaction as a virulence factor and major surface antigen of the blood stages of the parasite, responsible for maintaining chronic infection. Whilst important in the biology of P. falciparum, these genes (50 to 60 genes per parasite genome) are routinely excluded from whole genome analyses due to their hyper-diversity, achieved primarily through recombination. The PfEMP1 head structure almost always consists of a DBL-CIDR tandem. Categorised into different groups (upsA, upsB, upsC), different head structures have been associated with different ligand-binding affinities and disease severities. We study how conserved individual DBL types are at the country, regional, and local scales in Sub-Saharan Africa. Using publicly-available sequence datasets and a novel ups classification algorithm, cUps, we performed an in silico exploration of DBL conservation through time and space in Africa. In all three ups groups, the population structure of DBL types in Africa consists of variants occurring at rare, low, moderate, and high frequencies. Non-rare variants were found to be temporally stable in a local area in endemic Ghana. When inspected across different geographical scales, we report different levels of conservation; while some DBL types were consistently found in high frequencies in multiple African countries, others were conserved only locally, signifying local preservation of specific types. Underlying this population pattern is the composition of DBL types within each isolate DBL repertoire, revealed to also consist of a mix of types found at rare, low, moderate, and high frequencies in the population. We further discuss the adaptive forces and balancing selection, including host genetic factors, potentially shaping the evolution and diversity of DBL types in Africa.

genetics↗

A scalable method for identifying recombinants from unaligned sequences

Recombination is a fundamental process in molecular evolution, and the identification of recombinant sequences is of major interest for biologists. However, current methods for detecting recombinants only work for aligned sequences, often require a reference panel, and do not scale well to large datasets. Thus they are not suitable for the analyses of highly diverse genes, such as the var genes of the malaria parasite Plasmodium falciparum, which are known to diversify primarily through recombination. We introduce an algorithm to detect recombinant sequences from an unaligned dataset. Our approach can effectively handle thousands of sequences without the need of an alignment or a reference panel, offering a general tool suitable for the analysis of many different types of sequences. We demonstrate the effectiveness of our algorithm through extensive numerical simulations; in particular, it maintains its accuracy in the presence of insertions and deletions. We apply our algorithm to a dataset of 17,335 DBL types in var genes from Ghana, enabling the comparison between recombinant and non-recombinant types for the first time. We observe that sequences belonging to the same ups type or DBL subclass recombine amongst themselves more frequently, and that non-recombinant DBL types are more conserved than recombinant ones. Author summaryRecombination is a fundamental process in molecular evolution where two genes exchange genetic material, diversifying the genes. It is important to properly model this process when reconstructing evolutionary history, and to do so we need to be able to identify recombinant genes. In this manuscript, we develop a method for this which can be applied to scenarios where current methods often fail, such as where genes are very diverse. We specifically focus on detecting recombinants in the var genes of the malaria parasite Plasmodium falciparum. These genes influence the length and severity of malaria infection, and therefore their study is critical to the treatment and prevention of malaria. They are also highly diverse, primarily because of recombination. Our analysis of genes from a cross-sectional study in Ghana study show fundamental relations between the patterns and prevalence of recombination in these genes and other important biological categorisations.

evolutionary biology↗