bioRxiv ScienceSearch

Biology subjects

Schaefer, R. J.

Publications and source records attributed to Schaefer, R. J..

2 recordsLinked to original sources

Integrating co-expression networks with GWAS detects genes driving elemental accumulation in maize seeds

BackgroundGenome wide association studies (GWAS) have identified thousands of loci linked to hundreds of traits in many different species. However, because linkage equilibrium implicates a broad region surrounding each identified locus, the causal genes often remain unknown. This problem is especially pronounced in non-human, non-model species where functional annotations are sparse and there is frequently little information available for prioritizing candidate genes.\n\nResultsTo address this issue, we developed a computational approach called Camoco (Co-Analysis of Molecular Components) that systematically integrates loci identified by GWAS with gene co-expression networks to prioritize putative causal genes. We applied Camoco to prioritize candidate genes from a large-scale GWAS examining the accumulation of 17 different elements in maize seeds. Camoco identified statistically significant subnetworks for the majority of traits examined, producing a prioritized list of high-confidence causal genes for several agronomically important maize traits. Two candidate genes identified by our approach were validated through analysis of mutant phenotypes. Strikingly, we observed a strong dependence in the performance of our approach on the type of co-expression network used: expression variation across genetically diverse individuals in a relevant tissue context (in our case, maize roots) outperformed other alternatives.\n\nConclusionsOur study demonstrates that co-expression networks can provide a powerful basis for prioritizing candidate causal genes from GWAS loci, but suggests that the success of such strategies can highly depend on the gene expression data context. Both the Camoco software and the lessons on integrating GWAS data with co-expression networks generalize to species beyond maize.

systems biology

Development of a high-density, 2M SNP genotyping array and 670k SNP imputation array for the domestic horse

BackgroundTo date, genome-scale analyses in the domestic horse have been limited by suboptimal single nucleotide polymorphism (SNP) density and uneven genomic coverage of the current SNP genotyping arrays. The recent availability of whole genome sequences has created the opportunity to develop a next generation, high-density equine SNP array.\n\nResultsUsing whole genome sequence from 153 individuals representing 24 distinct breeds collated by the equine genomics community, we cataloged over 23 million de novo discovered genetic variants. Leveraging genotype data from individuals with both whole genome sequence, and genotypes from lower-density, legacy SNP arrays, a subset of [~]5 million high-quality, high-density array candidate SNPs were selected based on breed representation and uniform spacing across the genome. Considering probe design recommendations from a commercial vendor (Affymetrix, now Thermo Fisher Scientific) a set of [~]2 million SNPs were selected for a next-generation high-density SNP chip (MNEc2M). Genotype data were generated using the MNEc2M array from a cohort of 332 horses from 20 breeds and a lower-density array, consisting of [~]670 thousand SNPs (MNEc670k), was designed for genotype imputation.\n\nConclusionsHere, we document the steps taken to design both the MNEc2M and MNEc670k arrays, report genomic and technical properties of these genotyping platforms, and demonstrate the imputation capabilities of these tools for the domestic horse.

genomics