bioRxiv Science⌕ Search

Biology subjects

Laboulaye, R.

Publications and source records attributed to Laboulaye, R..

4 recordsLinked to original sources

ClOneHORT: Approaches for Improved Fidelity in Generative Models of Synthetic Genomes

MotivationDeep generative models have the potential to overcome difficulties in sharing individual-level genomic data by producing synthetic genomes that preserve the genomic associations specific to a cohort while not violating the privacy of any individual cohort member. However, there is significant room for improvement in the fidelity and usability of existing synthetic genome approaches. ResultsWe demonstrate that when combined with plentiful data and with population-specific selection criteria, deep generative models can produce synthetic genomes and cohorts that closely model the original populations. Our methods improve fidelity in the site-frequency spectra and linkage disequilibrium decay and yield synthetic genomes that can be substituted in downstream local ancestry inference analysis, recreating results with .91 to .94 accuracy. AvailabilityThe model described in this paper is freely available at github.com/rlaboulaye/clonehort.

genomics↗

Comparing the effect of imputation reference panel composition in four distinct Latin American cohorts

Genome-wide association studies have been useful in identifying genetic risk factors for various phenotypes. These studies rely on imputation and many existing panels are largely composed of individuals of European ancestry, resulting in lower levels of imputation quality in underrepresented populations. We aim to analyze how the composition of imputation reference panels affects imputation quality in four target Latin American cohorts. We compared imputation quality for chromosomes 7 and X when altering the imputation reference panel by: 1) increasing the number of Latin American individuals; 2) excluding either Latin American, African, or European individuals, or 3) increasing the Indigenous American (IA) admixture proportions of included Latin Americans. We found that increasing the number of Latin Americans in the reference panel improved imputation quality in the four populations; however, there were differences between chromosomes 7 and X in some cohorts. Excluding Latin Americans from analysis resulted in worse imputation quality in every cohort, while differential effects were seen when excluding Europeans and Africans between and within cohorts and between chromosomes 7 and X. Finally, increasing IA-like admixture proportions in the reference panel increased imputation quality at different levels in different populations. The difference in results between populations and chromosomes suggests that existing and future reference panels containing Latin American individuals are likely to perform differently in different Latin American populations.

genomics↗

Strong Positive Selection Biases Identity-By-Descent-Based Inferences of Recent Demography and Population Structure in Plasmodium falciparum

Malaria genomic surveillance often estimates parasite genetic relatedness using metrics such as Identity-By-Decent (IBD). Yet, strong positive selection stemming from antimalarial drug resistance or other interventions may bias IBD-based estimates. In this study, we utilized simulations, a true IBD inference algorithm, and empirical datasets from different malaria transmission settings to investigate the extent of such bias and explore potential correction strategies. We analyzed whole genome sequence data generated from 640 new and 4,026 publicly available Plasmodium falciparum clinical isolates. Our findings demonstrated that positive selection distorts IBD distributions, leading to underestimated effective population size and blurred population structure. Additionally, we discovered that the removal of IBD peak regions partially restored the accuracy of IBD-based inferences, with this effect contingent on the populations background genetic relatedness. Consequently, we advocate for selection correction for parasite populations undergoing strong, recent positive selection, particularly in high malaria transmission settings.

evolutionary biology↗

Genetics of Latin American Diversity (GLAD) Project: insights into population genetics and association studies in recently admixed groups in the Americas

Latin America is underrepresented in genetic studies, which can exacerbate disparities in personalized genomic medicine. However, genetic data of thousands of Latin Americans are already publicly available, but require a bureaucratic maze to navigate all the data access and consenting issues. We present the Genetics of Latin American Diversity (GLAD) Project, a platform that compiles genome-wide information of 54,077 Latin Americans from 39 studies representing 45 geographical regions. Through GLAD, we identified heterogeneous ancestry composition and recent gene-flow across the Americas. Also, we developed a simulated-annealing-based algorithm to match the genetic background of external samples to our database and share summary statistics without transferring individual-level data. Finally, we demonstrate the potential of GLAD as a critical resource for evaluating statistical genetic softwares in the presence of admixture. By making this resource available, we promote genomic research in Latin Americans and contribute to the promises of personalized medicine to more people.

genetics↗