bioRxiv Science⌕ Search

Biology subjects

Kuehn, C.

Publications and source records attributed to Kuehn, C..

3 recordsLinked to original sources

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

genomics↗

Polyadenylated RNA sequencing analysis helps establish a reliable catalog of circular RNAs - a bovine example.

The aim of this study was to compare the circular transcriptome of divergent tissues in order to understand: i) the presence of circular RNAs (circRNAs) that are not exonic circRNAs, i.e. originated from backsplicing involving known exons and, ii) the origin of artificial circRNA (artif_circRNA), i.e. circRNA not generated in-vivo. CircRNA identification is mostly an in-silico process, and the analysis of data from the BovReg project (https://www.bovreg.eu/) provided an opportunity to explore new ways to identify reliable circRNAs. By considering 117 tissue samples, we characterized 23,926 exonic circRNAs, 337 circRNAs from 273 introns (191 ciRNAs, 146 intron circles), 108 circRNAs from small non-coding genes and nearly 36.6K circRNAs classified as other_circRNAs. We suggested in-vivo copying of specific exonic circRNAs by an RNA-dependent RNA polymerase (RdRP) to explain the 20 identified circRNAs with reverse-complement exons. Furthermore, for 63 of those samples we analyzed in parallel data from total-RNAseq (ribosomal RNAs depleted prior to library preparation) with paired mRNAseq (library prepared with poly(A)-selected RNAs). The high number of circRNAs detected in mRNAseq, and the significant number of novel circRNAs, mainly other_circRNAs, led us to consider all circRNAs detected in mRNAseq as artificial. This study provided evidence that there were 189 false entries in the list of exonic circRNAs: 103 artif_circRNAs identified through comparison of total-RNAseq/mRNAseq using two circRNA tools, 26 probable artif_circRNAs, and 65 identified through deep annotation analysis. This study demonstrates the effectiveness of a panel of highly expressed exonic circRNAs (5-8%) in analyzing the diversity of the bovine circular transcriptome.

genomics↗

Improving the annotation of the cattle genome by annotating transcription start sites in a diverse set of tissues and populations using CAGE sequencing

Understanding the genomic control of tissue-specific gene expression and regulation can help to inform the application of genomic technologies in farm animal breeding programmes. The fine mapping of promoters (transcription start sites [TSS]) and enhancers (divergent amplifying segments of the genome local to TSS) in different populations of cattle across a wide diversity of tissues provides information to locate and understand the genomic drivers of breed- and tissue-specific phenotypes. To this aim we used Cap Analysis Gene Expression (CAGE) sequencing to define TSS and their co-expressed short-range enhancers (<1kb) in the ARS-UCD1.2_Btau5.0.1Y reference genome (1000bulls run9) and analysed tissue- and population specificity of expressed promoters. We identified 51,295 TSS and 2,328 TSS-Enhancer regions shared across the three populations (Holstein, Charolais x Holstein and Kinsella beef composite [KC]). In addition, we performed a comparative analysis of our cattle dataset with available data for seven other species to identify TSS and TSS-Enhancers that are specific to cattle. The CAGE dataset will be combined with other transcriptomic information for the same tissues generated in the BovReg project to create a new high-resolution map of transcript diversity across tissues and populations in cattle. Here we provide the CAGE dataset and annotation tracks for TSS and TSS Enhancers in the cattle genome. This new annotation information will improve our understanding of the drivers of gene expression and regulation in cattle and help to inform the application of genomic technologies in breeding programmes.

genomics↗