bioRxiv Science⌕ Search

Biology subjects

Boichard, D.

Publications and source records attributed to Boichard, D..

6 recordsLinked to original sources

Deep Generative Models for Discrete Genotype Simulation

Deep generative models open new avenues for simulating realistic genomic data while preserving privacy and addressing data accessibility constraints. While previous studies have primarily focused on generating gene expression or haplotype data, this study explores generating genotype data in both unconditioned and phenotype-conditioned settings, which is inherently more challenging due to the discrete nature of genotype data. In this work, we developed and evaluated commonly used generative models, including Variational Autoencoders (VAEs), Diffusion Models, and Generative Adversarial Networks (GANs), and proposed adaptation tailored to discrete genotype data. We conducted extensive experiments on large-scale datasets, including all chromosomes from cow and multiple chromosomes from human. Model performance was assessed using a well-established set of metrics drawn from both deep learning and quantitative genetics literature. Our results show that these models can effectively capture genetic patterns and preserve genotype-phenotype association. Our findings provide a comprehensive comparison of these models and offer practical guidelines for future research in genotype simulation. We have made our code publicly available at https://github.com/SihanXXX/DiscreteGenoGen.

bioinformatics↗

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

genomics↗

Comprehensive detection of structural variations in long and short reads dataset of French cattle

Structural variants (SVs) correspond to different types of genomic variants larger than 50 bp. Many findings suggest the use of long rather than short reads to improve the accuracy of SV detection. Here, we present the results of an in-depth analysis for detection of SVs, mainly large insertions and deletions, in 14 French bovine breeds, based on whole-genome data comprising 176 long-read and 571 short-read samples, with 154 individuals having both long- and short-read data available. We first investigated possible biases on the performances of well-known SV detection tools, namely CUTESV, PBSV, and SNIFFLES, using long reads from different technologies, including PacBio HiFi, Oxford ONT, and PacBio CLR. We subsequently highlighted the abilities of tools for detecting SVs (DELLY, LUMPY, and MANTA) and for genotyping known SVs (GRAPHTYPER, SVTYPER, PARAGRAPH, and VG toolkit) using short-read data. We then show how the incremental composition of samples in the reference panel affected the SV genotyping for six validation individuals sequenced in short reads. We then searched for the optimal parameters and created the final SV reference panel consisting of 25,191 deletions and 30,118 insertions. Finally, we emphasized the landscape of the genotyped SVs segregating across 571 short-read individuals of 14 breeds.

genomics↗

Application of a French cattle pangenome, from structural variant discovery to association studies on key phenotypes

BackgroundThe current cattle reference genome assembly, a pseudo-linear sequence produced using sequences from a single Hereford cow, represent a limit when performing genetic studies, especially when investigating the whole spectrum of genetic variations within the species. Detecting structural variations (SVs) poses significant challenges when relying solely on conventional methods of short or long-read sequence mapping to the current bovine genome assembly. ResultsIn this study, we used long-reads (LR) and bioinformatic tools to construct a comprehensive bovine pangenome incorporating genetic diversity of 64 good quality de novo genome assemblies representing 14 French dairy and beef cattle breeds. Using a combination of complementary approaches, we explored the pangenome graph and identified 2.563 Gb of sequences common to all samples, and cumulated 0.295 Gb of variable sequences. Notably, we discovered 0.159 Gb of novel sequences not present in the current Hereford reference genome assembly. Our analysis also revealed 109,275 SVs, of which 84,612 were bi-allelic, including 21,840 insertions and 21,340 deletions. Genome-wide association studies using SNPs and a panel of 221 SVs, shared between the pangenome and the EuroGMD chip, revealed several well-known QTLs across the genome for the Holstein, Montbeliarde and Normande breeds. Among those, a QTL on chromosome 11 presents an SV with a highly significant effect on stature in the Holstein breed. This SV is a 6.2 kb deletion affecting the 5UTR, first exon and part of first intron of MATN3 gene, suggesting a potential regulatory and coding effect. ConclusionsOur study provides new insights into the genetic diversity of 14 French dairy and beef breeds and highlights the utility of pangenome graphs in capturing structural variation. The identified SV associated with stature highlights the importance of integrating SVs into GWAS for a more comprehensive understanding of complex traits.

genetics↗

Trajectories of genetic correlations in populations under selection: from theory to a case-study

BackgroundBreeding programs select for multiple commercial traits, aiming to achieve genetic progress for all. Often, selection is based on a selection index, i.e. a linear combination of traits with weights defined by, among other information, the genetic correlation between traits. These correlations are typically estimated as a static parameter, and assumed equal to all individuals and generations. While research on the consequences of selection to genetic variances (Bulmer effect) is widely available, only a few studies focused on the consequences of selection to genetic correlations. Our study extended the already existing inferences about how selection affects genetic variances, to how multi-trait selection affects genetic correlations. In order to further our understanding of genetic correlations, we also proposed an alternative method to calculate genetic correlations between traits at the individual level, called by us as individualized sire genetic correlation (iSGC), obtained through the estimated breeding values (EBV) from evaluated daughters. Lastly, a case-study was performed on thirty years of data from the French Holstein dairy cattle population, for five traits studied pairwise: milk and protein yield, milking speed, somatic cell score, and cow conception rate. ResultsTheory revealed that multi-trait selection leads to an attenuation (decrease) of positive genetic correlations, with potential to revert them to negative values, if initially low. Uncorrelated traits will become negatively correlated, and negative genetic correlations will be either intensified or attenuated (decrease or increase, respectively), depending on selection intensity, weights applied to the selection index, and the initial genetic correlation. ConclusionBoth theory and empirical results on real data confirm that selection does change the genetic correlation between traits in a population under selection. Moreover, empirical trajectories of the iSGC were in better agreement with the theory, than trajectories of populational genetic correlations. The iSGC searches for individual-specific patterns of correlations, and since it is measured on sires through the EBV of their daughters, it also considers the recombination of the genetic background. Along with the fact that trajectories of iSGC were in better agreement with theory, we believe it to be a potentially less biased measure of genetic correlations between traits.

genetics↗

A bovine model of rhizomelic chondrodysplasia punctata caused by a deep intronic splicing mutation in the GNPAT gene

BackgroundGenetic defects that occur naturally in livestock species provide valuable models for investigating the molecular mechanisms underlying rare human diseases. Livestock breeds are subject to the regular emergence of recessive genetic defects, due to their low genetic variability, while their large population sizes provide easy access to case and control individuals, as well as massive amounts of pedigree, genomic and phenotypic information recorded for selection purposes. In this study, we investigated a lethal form of recessive chondrodysplasia observed in 21 stillborn calves of the Aubrac breed of beef cattle. ResultsDetailed clinical examinations revealed proximal limb shortening, epiphyseal calcific deposits and other clinical signs consistent with human rhizomelic chondrodysplasia punctata, a rare peroxisomal disorder caused by recessive mutations in one of five genes (AGPS, FAR1, GNPAT, PEX5 and PEX7). Using homozygosity mapping, whole genome sequencing of two affected individuals, and filtering for variants found in 1,867 control genomes, we reduced the list of candidate variants to a single deep intronic substitution in GNPAT (g.4,039,268G>A on Chromosome 28 of the ARS-UCD1.2 bovine genome assembly). For verification, we performed large-scale genotyping of this variant using a custom SNP array and found a perfect genotype-phenotype correlation in 21 cases and 26 of their parents, and a complete absence of homozygotes in 1,195 Aubrac controls. The g.4,039,268A allele segregated at a frequency of 2.6% in this population and was absent in 375,535 additional individuals from 17 breeds. Then, using in vivo and in vitro analyses, we demonstrated that the derived allele activates cryptic splice sites within intron 11 resulting in abnormal transcripts. Finally, by mining the wealth of records available in the French bovine database, we demonstrated that this deep intronic substitution was responsible not only for stillbirth but also for juvenile mortality in homozygotes and had a moderate but significant negative effect on muscle development in heterozygotes. ConclusionsWe report the first spontaneous large animal model of rhizomelic chondrodysplasia punctata and provide both a diagnostic test to counter-select this defect in cattle and interesting insights into the molecular consequences of complete or partial GNPAT insufficiency in mammals.

genetics↗