bioRxiv Science⌕ Search

Biology subjects

Marsh, J. I.

Publications and source records attributed to Marsh, J. I..

4 recordsLinked to original sources

Biases in ARG-based inference of historical population size in populations experiencing selection

Inferring the demographic history of populations provides fundamental insights into species dynamics and is essential for developing a null model to accurately study selective processes. However, background selection and selective sweeps can produce genomic signatures at linked sites that mimic or mask signals associated with historical population size change. While the theoretical biases introduced by the linked effects of selection have been well established, it is unclear whether ARG-based approaches to demographic inference in typical empirical analyses are susceptible to mis-inference due to these effects. To address this, we developed highly realistic forward simulations of human and Drosophila melanogaster populations, including empirically estimated variability of gene density, mutation rates, recombination rates, purifying and positive selection, across different historical demographic scenarios, to broadly assess the impact of selection on demographic inference using a genealogy-based approach. Our results indicate that the linked effects of selection minimally impact demographic inference for human populations, though it could cause mis-inference in populations with similar genome architecture and population parameters experiencing more frequent recurrent sweeps. We found that accurate demographic inference of D. melanogaster populations by ARG-based methods is compromised by the presence of pervasive background selection alone, leading to spurious inferences of recent population expansion which may be further worsened by recurrent sweeps, depending on the proportion and strength of beneficial mutations. Caution and additional testing with species-specific simulations are needed when inferring population history with non-human populations using ARG-based approaches to avoid mis-inference due to the linked effects of selection.

genetics↗

Adaptive gene loss in the common bean pan-genome during range expansion and domestication

The common bean (Phaseolus vulgaris L.) is a crucial grain legume crop [1,2] whose life history offers an ideal evolutionary model to identify and study adaptive variants in wild and domestication populations [3]. Here we present the first common bean pan-genome based on five high-quality genomes and whole-genome reads representing 339 genotypes. We found [~]243 Mb of additional sequences containing 7,495 protein-coding genes missing from the reference, constituting 51% of the total presence/absence variations (PAVs). There were more putatively deleterious mutations in PAVs than core genes, probably reflecting the lower effective population size of PAVs as well as fitness advantages due to the purging effect of gene loss. Our results suggest strong pan-genome shrinkage occurred during wild range expansion from Mexico to South America, with more PAV loss per individual in Andean vs Mesoamerican populations. Selection signatures during wild spreading and domestication were also associated with PAV loss involved in important adaptive traits. Our findings provide evidence that partial or complete gene loss was a key adaptive trait leading to localized and genome-wide reductions. This novel result has major implications for the understanding of the process of plant adaptation and claims for a paradigm shift in evolutionary genetics. Moreover, the common bean pan-genome is a valuable resource for food legume research and breeding towards climate change mitigation, and sustainable agriculture.

evolutionary biology↗

Local haplotype visualization for trait association analysis with crosshap

SummaryGWAS excels at harnessing dense genomic variant datasets to identify candidate regions responsible for producing a given phenotype. However, GWAS and traditional fine-mapping methods do not provide insight into the complex local landscape of linkage that contains and has been shaped by the causal variant(s). Here, we present crosshap, an R package that performs robust density-based clustering of variants based on their linkage profiles to capture haplotype structures in a local genomic region of interest. Following this, crosshap is equipped with visualization tools for choosing optimal clustering parameters ({varepsilon}) before producing an intuitive figure that provides an overview of the complex relationships between linked variants, haplotype combinations, phenotypic traits and metadata. Availability and implementationThe crosshap package is freely available under the MIT license and can be downloaded directly from CRAN with R>4.0.0. The development version is available on GitHub alongside issue support (https://github.com/jacobimarsh/crosshap). Tutorial vignettes and documentation are available (https://jacobimarsh.github.io/crosshap/).

bioinformatics↗

Haplotype mapping uncovers unexplored variation in wild and domesticated soybean at the major protein locus cqProt-003

Here, we present association and linkage analysis of 985 wild, landrace and cultivar soybean accessions in a pan genomic dataset to characterize the major high-protein/low-oil associated locus cqProt-003 located on chromosome 20. A significant trait associated region within a 173 kb linkage block was identified and variants in the region were characterised, identifying 34 high confidence SNPs, 4 insertions, 1 deletion and a larger 304 bp structural variant in the high-protein haplotype. Trinucleotide tandem repeats of variable length present in the third exon of gene 20G085100 are strongly correlated with the high-protein phenotype and likely represent causal variation. Structural variation has previously been found in the same gene, for which we report the global distribution of the 304bp deletion and have identified additional nested variation present in high-protein individuals. Mapping variation at the cqProt-003 locus across demographic groups suggests that the high-protein haplotype is common in wild accessions (94.7%), rare in landraces (10.6%) and near absent in cultivated breeding pools (4.1%), suggesting its decrease in frequency primarily correlates with domestication and continued during subsequent improvement. However, the variation that has persisted in under-utilized wild and landrace populations holds high breeding potential for breeders willing to forego seed oil to maximise protein content. The results of this study include the identification of distinct haplotype structures within the high-protein population, and a broad characterization of the genomic context and linkage patterns of cqProt-003 across global populations, supporting future functional characterisation and modification. Key messageThe major soy protein QTL, cqProt-003, was analysed for haplotype diversity and global distribution, results indicate 304bp deletion and variable tandem repeats in protein coding regions are likely causal candidates.

plant biology↗