bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Genome-culture coevolution promotes rapid divergence in the killer whale

The interaction between ecology, culture and genome evolution remains poorly understood. Analysing population genomic data from killer whale ecotypes, which we estimate have globally radiated within less than 250,000 years, we show that genetic structuring including the segregation of potentially functional alleles is associated with socially inherited ecological niche. Reconstruction of ancestral demographic history revealed bottlenecks during founder events, likely promoting ecological divergence and genetic drift resulting in a wide range of genome-wide differentiation between pairs of allopatric and sympatric ecotypes. Functional enrichment analyses provided evidence for regional genomic divergence associated with habitat, dietary preferences and postzygotic reproductive isolation. Our findings are consistent with expansion of small founder groups into novel niches by an initial plastic behavioural response, perpetuated by social learning imposing an altered natural selection regime. The study constitutes an important step toward an understanding of the complex interaction between demographic history, culture, ecological adaptation and evolution at the genomic level.

Genomics

Integrating genomics into clinical pediatric oncology using the molecular tumor board at the Memorial Sloan Kettering Cancer Center

BackgroundPediatric oncologists have begun to leverage tumor genetic profiling to match patients with targeted therapies. At the Memorial Sloan Kettering Cancer Center (MSKCC), we developed the Pediatric Molecular Tumor Board (PMTB) to track, integrate, and interpret clinical genomic profiling and potential targeted therapeutic recommendations.\n\nProcedureThis retrospective case series includes all patients reviewed by the MSKCC PMTB from July 2014 to June 2015. Cases were submitted by treating oncologists and potential treatment recommendations were based upon the modified guidelines of the Oxford Centre for Evidence Based Medicine.\n\nResultsThere were 41 presentations of 39 individual patients during the study period. Gliomas, acute myeloid leukemia, and neuroblastoma were the most commonly reviewed cases. Thirty nine (87%) of the 45 molecular sequencing profiles utilized hybrid-capture targeted genome sequencing. In 30 (73%) of the 41 presentations, the PMTB provided therapeutic recommendations, of which 19 (46%) were implemented. Twenty-one (70%) of the recommendations involved targeted therapies. Three (14%) targeted therapy recommendations had published evidence to support the proposed recommendations (evidence levels 1-2), 8 (36%) recommendations had preclinical evidence (level 3), and 11 (50%) recommendations were based upon hypothetical biological rationales (level 4).\n\nConclusionsThe MSKCC PMTB enabled a clinically relevant interpretation of genomic profiling. Effective use of clinical genomics is anticipated to require new and improved tools to ascribe pathogenic significance and therapeutic actionability. Development of specific rule-driven clinical protocols will be needed for the incorporation and evaluation of genomic and molecular profiling in interventional prospective clinical trials.

Genomics

Whole-genome characterization in pedigreed non-human primates using Genotyping-By-Sequencing and imputation.

BackgroundRhesus macaques are widely used in biomedical research, but the application of genomic information in this species to better understand human disease is still undeveloped. Whole-genome sequence (WGS) data in pedigreed macaque colonies could provide substantial experimental power, but the collection of WGS data in large cohorts remains a formidable expense. Here, we describe a cost-effective approach that selects the most informative macaques in a pedigree for whole-genome sequencing, and imputes these dense marker data into all remaining individuals having sparse marker data, obtained using Genotyping-By-Sequencing (GBS).\n\nResultsWe developed GBS for the macaque genome using a single digest with PstI, followed by sequencing to 30X coverage. From GBS sequence data collected on all individuals in a 16-member pedigree, we characterized an optimal 22,455 sparse markers spaced ~125 kb apart. To characterize dense markers for imputation, we performed WGS at 30X coverage on 9 of the 16 individuals, yielding ~10.2 million high-confidence variants. Using the approach of \"Genotype Imputation Given Inheritance\" (GIGI), we imputed alleles at an optimized dense set of 4,920 variants on chromosome 19, using 490 sparse markers from GBS. We assessed changes in accuracy of imputed alleles, 1) across 3 different strategies for selecting individuals for WGS, i.e., a) using \"GIGI-Pick\" to select informative individuals, b) sequencing the most recent generation, or c) sequencing founders only; and 2) when using from 1-9 WGS individuals for imputation. We found that accuracy of imputed alleles was highest using the GIGI-Pick selection strategy (median 92%), and improved very little when using >4 individuals with WGS for imputation. We used this ratio of 4 WGS to 12 GBS individuals to impute an expanded set of ~14.4 million variants across all 20 macaque autosomes, achieving ~85-88% accuracy per chromosome.\n\nConclusionsWe conclude that an optimal tradeoff exists at the ratio of 1 individual selected for WGS using the GIGI-Pick algorithm, per 3-5 relatives selected for GBS, a cost savings of ~67-83% over WGS of all individuals. This approach makes feasible the collection of accurate, dense genome-wide sequence data in large pedigreed macaque cohorts without the need for expensive WGS data on all individuals.

Genomics

Wide genome involvement in response to long-term selection for antibody response in an experimental population of White Leghorn chickens

Long-term selection experiments provide a powerful approach to gain empirical insights into adaptation. They allow researchers to uncover the targets of selection and how these contribute to the mode and tempo of adaptation. Here we report results from a pooled genome re-sequencing study to investigate the consequences of 39 generations of bidirectional selection in White Leghorn chickens on a humoral immune trait: antibody response to sheep red blood cells. We observed wide genome involvement in response to this selection regime, with over 200 candidate sweep regions characterised by spans of high genetic differentiation (FST). These sweep signatures, encompassing almost 20% of the chicken genome (208.8 Mb), are primarily the result from bidirectional selection on haplotypes present in the base population. These extensive genomic changes highlight both the extent of standing genetic variation at immune loci available at the onset of selection, as well as how the long-term selection response results from selection on a highly polygenic genetic architecture. Furthermore, we present three examples of strong candidate genes that may have contributed to the profound phenotypic response to selection.\n\nData AvailabilityPooled genome data generated for this study will become available via SRA upon acceptance of manuscript

Genomics

Whole genome SNP typing to investigate methicillin-resistant Staphylococcus aureus carriage in a health-care provider as the source of multiple surgical site infections.

BackgroundPrevention of nosocomial transmission of infections is a central responsibility in the healthcare environment, and accurate identification of transmission events presents the first challenge. Phylogenetic analysis based on whole genome sequencing provides a high-resolution approach for accurately relating isolates to one another, allowing precise identification or exclusion of transmission events and sources for nearly all cases. We sequenced 24 methicillin-resistant Staphylococcus aureus (MRSA) genomes to retrospectively investigate a suspected point source of three surgical site infections (SSIs) that occurred over a one-year period. The source of transmission was believed to be a surgical team member colonized with MRSA, involved in all surgeries preceding the SSI cases, who was subsequently decolonized. Genetic relatedness among isolates was determined using whole genome single nucleotide polymorphism (SNP) data.\n\nResultsWhole genome SNP typing (WGST) revealed 283 informative SNPs between the surgical team members isolate and the closest SSI isolate. The second isolate was 286 and the third was thousands of SNPs different, indicating the nasal carriage strain from the surgical team member was not the source of the SSIs. Given the mutation rates estimated for S. aureus, none of the SSI isolates share a common ancestor within the past 14 years, further discounting any common point source for these infections. The decolonization procedures and resources spent on the point source infection control could have been prevented if WGST was performed at the time of the suspected transmission, instead of retrospectively.\n\nConclusionsWhole genome sequence analysis is an ideal method to exclude isolates involved in transmission events and nosocomial outbreaks, and coupling this method with epidemiological data can determine if a transmission event occurred. These methods promise to direct infection control resources more appropriately.

Genomics

Genomic Prediction of Hybrid Combinations in the Early Stages of a Maize Hybrid Breeding Pipeline

Prediction of single-cross hybrid performance has been a major goal of plant breeders since the beginning of hybrid breeding. Genomic prediction has shown to be a promising approach, but only limited studies have examined the accuracy of predicting single cross performance. Most of the studies rather focused on predicting top cross performance using single tester to determine the inbred parents worth in hybrid combinations. Moreover, no studies have examined the potential of predicting single crosses made among random progenies derived from a series of biparental families, which resembles the structure of germplasm comprising the initial stages of a hybrid maize breeding pipeline. The main objective of this study was to evaluate the potential of genomic prediction for identifying superior single crosses early in the breeding pipeline and optimize its application. To accomplish these objectives, we designed and analyzed a novel population of single-cross hybrids representing the Iowa Stiff Stalk Synthetic/Non-Stiff Stalk heterotic pattern commonly used in the development of North American commercial maize hybrids. The single cross prediction accuracies estimated using cross-validation ranged from 0.40 to 0.74 for grain yield, 0.68 to 0.91 for plant height and 0.54 to 0.94 for staygreen depending on the number of tested parents of the single crosses. The genomic estimated general and specific combining abilities showed a clear advantage over the use of genomic covariances among single crosses, especially when one or both parents of the single cross were untested in hybrid combinations. Overall, our results suggest that genomic prediction of the performance of single crosses made using random progenies from the early stages of the breeding pipeline holds great potential to re-design hybrid breeding and increase its efficiency.

Genomics

Forward genetic screen of human transposase genomic rearrangements

Background. Numerous human genes encode potentially active DNA transposases or recombinases, but our understanding of their functions remains limited due to shortage of methods to profile their activities on endogenous genomic substrates. Results. To enable functional analysis of human transposase-derived genes, we combined forward chemical genetic hypoxanthine-guanine phosphoribosyltransferase 1 (HPRT1) screening with massively parallel paired-end DNA sequencing and structural variant genome assembly and analysis. Here, we report the HPRT1 mutational spectrum induced by the human transposase PGBD5, including PGBD5-specific signal sequences (PSS) that serve as potential genomic rearrangement substrates. Conclusions. The discovered PSS motifs and high-throughput forward chemical genomic screening approach should prove useful for the elucidation of endogenous genome remodeling activities of PGBD5 and other domesticated human DNA transposases and recombinases.

Genomics

Regions of very low H3K27me3 partition the Drosophila genome into topological domains

BackgroundIt is now well established that eukaryote genomes have a common architectural organization into topologically associated domains (TADs) and evidence is accumulating that this organization plays an important role in gene regulation. However, the mechanisms that partition the genome into TADs and the nature of domain boundaries are still poorly understood.\n\nResultsWe have investigated boundary regions in the Drosophila genome and find that they can be identified as domains of very low H3K27me3. The genome-wide H3K27me3 profile partitions into two states; very low H3K27me3 identifies Depleted (D) domains that contain housekeeping genes and their regulators such as the histone acetyltransferase-containing NSL complex, whereas domains containing mid-to-high levels of H3K27me3 (Enriched or E domains) are associated with regulated genes, irrespective of whether they are active or inactive. The D domains correlate with the boundaries of TADs and are enriched in a subset of architectural proteins, particularly Chromator, BEAF-32, and Z4/Putzig. However, rather than being clustered at the borders of these domains, these proteins bind throughout the H3K27me3-depleted regions and are much more strongly associated with the transcription start sites of housekeeping genes than with the H3K27me3 domain boundaries.\n\nConclusionsWe suggest that the D domain chromatin state, characterised by very low H3K27me3 and established by housekeeping gene regulators, acts to separate topological domains thereby setting up the domain architecture of the genome.

Genomics

Chironomus riparius (Diptera) genome sequencing reveals the impact of minisatellite transposable elements on population divergence

Active transposable elements (TEs) may result in divergent genomic insertion and abundance patterns among conspecific populations. Upon secondary contact, such divergent genetic backgrounds can theoretically give rise to classical Dobzhansky-Muller incompatibilities (DMI), a way how TEs can contribute to the evolution of endogenous genetic barriers and eventually population divergence. We investigated whether differential TE activity created endogenous selection pressures among conspecific populations of the non-biting midge Chironomus riparius, focussing on a Chironomus-specific TE, the minisatellite-like Cla-element, whose activity is associated with speciation in the genus. Using an improved and annotated draft genome for a genomic study with five natural C. riparius populations, we found highly population-specific TE insertion patterns with many private insertions. A highly significant correlation of pairwise population FST from genome-wide SNPs with the FST estimated from TEs suggests drift as the major force driving TE population differentiation. However, the significantly higher Cla-element FST level due to a high proportion of differentially fixed Cla-element insertions indicates that segregating, i.e. heterozygous insertions are selected against. With reciprocal crossing experiments and fluorescent in-situ hybridisation of Cla-elements to polytene chromosomes, we documented phenotypic effects on female fertility and chromosomal mispairings that might be linked to DMI in hybrids. We propose that the inferred negative selection on heterozygous Cla-element insertions causes endogenous genetic barriers and therefore acts as DMI among C. riparius populations. The intrinsic genomic turnover exerted by TEs, thus, may have a direct impact on population divergence that is operationally different from drift and local adaptation.

genomics

Chromosomal dynamics predicted by an elastic network model explains genome-wide accessibility and long-range couplings

Understanding the three-dimensional (3D) architecture of the chromatin and its relation to gene expression and regulation is fundamental to understanding how the genome functions. Advances in Hi-C technology now permit us to have a glimpse into the 3D genome organization and identify topologically associated domains (TADs), but we still lack an understanding of the structural dynamics of chromosomes. The dynamic couplings between regions separated by large genomic distances (> 50 megabases) have yet to be characterized. We adapted a well-established protein-modeling framework, the Gaussian Network Model (GNM), to the task of modeling chromatin dynamics using Hi-C contact data. We show that the GNM can identify structural dynamics at multiple scales: it can quantify the fluctuations in the positions of gene loci, find large genomic compartments and smaller TADs that undergo en-bloc movements, and identify dynamically coupled distal regions along the chromosomes. We show that the predictions of the GNM correlate well with DNase-seq and ATAC-seq measurements on accessibility, the previously identified A and B compartments of chromatin structure, and pairs of interacting loci identified by ChIA-PET. We describe a method to use the GNM to identify novel cross-correlated distal domains (CCDDs) representing regions of long-range dynamic coupling and show that CCDDs are often associated with increased gene coexpression using a large-scale analysis of 212 expression experiments. Together, these results show that GNM provides a mathematically well-founded unified framework for assessing chromatin dynamics and the structural basis of genome-wide observations.

genomics

Whole genome resequencing of a laboratory-adapted Drosophila melanogaster population sample

As part of a study into the molecular genetics of sexually dimorphic complex traits, we used next-generation sequencing to obtain data on genomic variation in an outbred laboratory-adapted fruit fly (Drosophila melanogaster) population. We successfully resequenced the whole genome of 2 females from the Berkeley reference line (BDGP6/dm6), and 220 hemiclonal females that were heterozygous for the same reference line genome, and a unique haplotype from the outbred base population (LHM). The use of a static and known genetic background enabled us to obtain sequences from whole-genome phased haplotypes. We used a BWA-Picard-GATK pipeline for mapping sequence reads to the dm6 reference genome assembly, at a median depth-of coverage of 31X, and have made the resulting data publicly-available in the NCBI Short Read Archive (BioProject PRJNA282591). Haplotype Caller discovered and genotyped 1,726,931 genetic variants (SNPs and indels, <200bp). Additionally, we used GenomeStrip/2.0 to discover and genotype 167 large structural variants (1-100Kb in size). Sequence data and quality-filtered genotype data are publicly-available at NCBI (Short Read Archive, dbSNP and dbVar). We have also released the unfiltered genotype data, and the code and logs for data processing, summary statistics, and graphs, via the research data repository, Zenodo, (https://zenodo.org/, Sussex Drosophila Sequencing community).

genomics

The megabase-sized fungal genome of Rhizoctonia solani assembled from nanopore reads only.

The ability to quickly obtain accurate genome sequences of eukaryotic pathogens at low costs provides a tremendous opportunity to identify novel targets for therapeutics, develop pesticides with increased target specificity and breed for resistance in food crops. Here, we present the first report of the ~54 MB eukaryotic genome sequence of Rhizoctonia solani, an important pathogenic fungal species of maize, using nanopore technology. Moreover, we show that optimizing the strategy for wet-lab procedures aimed to isolate high quality and ultra-pure high molecular weight (HMW) DNA results in increased read length distribution and thereby allowing generation of the most contiguous genome assembly for R. solani to date. We further determined sequencing accuracy and compared the assembly to short-read technologies. With the current sequencing technology and bioinformatics tool set, we are able to deliver an eukaryotic fungal genome at low cost within a week. With further improvements of the sequencing technology and increased throughput of the PromethION sequencer we aim to generate near-finished assemblies of large and repetitive plant genomes and cost-efficiently perform de novo sequencing of large collections of microbial pathogens and the microbial communities that surround our crops.

genomics

Complete avian malaria parasite genomes reveal host-specific parasite evolution in birds and mammals

Avian malaria parasites are prevalent around the world, and infect a wide diversity of bird species. Here we report the sequencing and analysis of high quality draft genome sequences for two avian malaria species, Plasmodium relictum and Plasmodium gallinaceum. We identify 50 genes that are specific to avian malaria, located in an otherwise conserved core of the genome that shares gene synteny with all other sequenced malaria genomes. Phylogenetic analysis suggests that the avian malaria species form an outgroup to the mammalian Plasmodium species and using amino acid divergence between species, we estimate the avian and mammalian-infective lineages diverged in the order of 10 million years ago. Consistent with their phylogenetic position, we identify orthologs of genes that had previously appeared to be restricted to the clades of parasites containing P. falciparum and P. vivax - the species with the greatest impact on human health. From these orthologs, we explore differential diversifying selection across the genus and show that the avian lineage is remarkable in the extent to which invasion related genes are evolving. The subtelomeres of the P. relictum and P. gallinaceum genomes contain several novel gene families, including an expanded surf multigene family. We also identify an expansion of reticulocyte binding protein homologs in P. relictum and within these proteins, we detect distinct regions that are specific to non-human primate, humans, rodent and avian hosts. For the first time in the Plasmodium lineage we find evidence of transposable elements, including several hundred fragments of LTR-retrotransposons in both species and an apparently complete LTR-retrotransposon in the genome of P. gallinaceum.

genomics

An annotated draft genome for Radix auricularia(Gastropoda, Mollusca)

Molluscs are the second most species-rich phylum in the animal kingdom, yet only eleven genomes of this group have been published so far. Here, we present the draft genome sequence of the pulmonate freshwater snail Radix auricularia. Six whole genome shotgun libraries with different layouts were sequenced. The resulting assembly comprises 4,823 scaffolds with a cumulative length of 910 Mb and an overall read coverage of 72x. The assembly contains 94.6 % of a metazoan core gene collection, indicating an almost complete coverage of the coding fraction. The discrepancy of ~690 Mb compared to the estimated genome size of R. auricularia (1.6 Gb) results from a high repeat content of 70 % mainly comprising DNA transposons. The annotation of 17,338 protein coding genes was supported by the use of publicly-available transcriptome data. This draft will serve as starting point for further genomic and population genetic research in this scientifically important phylum.

genomics

Recombinant DNA resources for the comparative genomics of Ancylostoma ceylanicum.

We describe the construction and initial characterization of genomic resources (a set of recombinant DNA libraries, representing in total over 90,000 independent plasmid clones), originating from the genome of a hamster adapted hookworm, Ancylostoma ceylanicum. First, with the improved methodology, we generated sets of SL1 (5 -linker - GGTTAATTACCCAAGTTTGAG), and captured cDNAs from two different hookworm developmental stages: pre-infective L3 and parasitic adults. Second, we constructed a small insert (2-10kb) genomic library. Third, we generated a Bacterial Artificial Chromosome library (30-60kb). To evaluate the quality of our libraries we characterized sequence tags on randomly chosen clones and with first pass screening we generated almost a hundred novel hookworm sequence tags. The sequence tags detected two broad classes of genes: i. conserved nematode genes and ii. putative hookworm-specific proteins. Importantly, some of the identified genes encode proteins of general interest including potential targets for hookworm control. Additionally, we identified a syntenic region in the mitochondrial genome, where the gene order is shared between the free-living nematode C. elegans and A. ceylanicum. Our results validate the use of recombinant DNA resources for comparative genomics of nematodes, including the free-living genetic model organism C. elegans and closely related parasitic species. We discuss the potential and relevance of Ancylostoma ceylanicum data and resources generated by the recombinant DNA approach.

genomics

10 Simple Rules for Sharing Human Genomic Data

Introduction Introduction Conclusion Competing Interests References Delivery of the promise of precision medicine relies heavily on human genomic data sharing. Sharing genome data generated through publicly funded projects maximises return on investment from taxpayer funds and increases the likelihood of obtaining funding in future rounds [1]. More importantly, genome data sharing makes it possible for other scientists to reuse existing datasets for further research and constitutes a direct measure of the current advancement in risk prediction, diagnosis, and treatment for genomic disorders [2].\n\nSharing of human genomic data carries responsibilities to protect confidentiality and the privacy of research participants [3]. In certain cases data sharing may be complicated or limited by agreements with ...

genomics

Genome-wide analysis of ivermectin response by Onchocerca volvulus reveals that genetic drift and soft selective sweeps contribute to loss of drug sensitivity

BackgroundTreatment of onchocerciasis using mass ivermectin administration has reduced morbidity and transmission throughout Africa and Central/South America. Mass drug administration is likely to exert selection pressure on parasites, and phenotypic and genetic changes in several Onchocerca volvulus populations from Cameroon and Ghana - exposed to more than a decade of regular ivermectin treatment - have raised concern that sub-optimal responses to ivermectins anti-fecundity effect are becoming more frequent and may spread.\n\nMethodology/Principal FindingsPooled next generation sequencing (Pool-seq) was used to characterise genetic diversity within and between 108 adult female worms differing in ivermectin treatment history and response. Genome-wide analyses revealed genetic variation that significantly differentiated good responder (GR) and sub-optimal responder (SOR) parasites. These variants were not randomly distributed but clustered in ~31 quantitative trait loci (QTLs), with little overlap in putative QTL position and gene content between countries. Published candidate ivermectin SOR genes were largely absent in these regions; QTLs differentiating GR and SOR worms were enriched for genes in molecular pathways associated with neurotransmission, development, and stress responses. Finally, single worm genotyping demonstrated that geographic isolation and genetic change over time (in the presence of drug exposure) had a significantly greater role in shaping genetic diversity than the evolution of SOR.\n\nConclusions/SignificanceThis study is one of the first genome-wide association analyses in a parasitic nematode, and provides insight into the genomics of ivermectin response and population structure of O. volvulus. We argue that ivermectin response is a polygenically-determined quantitative trait in which identical or related molecular pathways but not necessarily individual genes likely determine the extent of ivermectin response in different parasite populations. Furthermore, we propose that genetic drift rather than genetic selection of SOR is the underlying driver of population differentiation, which has significant implications for the emergence and potential spread of SOR within and between these parasite populations.\n\nAuthor summaryOnchocerciasis is a human parasitic disease endemic across large areas of Sub-Saharan Africa, where more that 99% of the estimated 100 million people globally at-risk live. The microfilarial stage of Onchocerca volvulus causes pathologies ranging from mild itching to visual impairment and ultimately, irreversible blindness. Mass administration of ivermectin kills microfilariae and has an anti-fecundity effect on adult worms by temporarily inhibiting the development in utero and/or release into the skin of new microfilariae, thereby reducing morbidity and transmission. Phenotypic and genetic changes in some parasite populations that have undergone multiple ivermectin treatments in Cameroon and Ghana have raised concern that sub-optimal response to ivermectins anti-fecundity effect may increase in frequency, reducing the impact of ivermectin-based control measures. We used next generation sequencing of small pools of parasites to define genome-wide genetic differences between phenotypically characterised good and sub-optimal responder parasites from Cameroon and Ghana, and identified multiple genomic regions differentiating the response types. These regions were largely different between parasites from both countries but revealed common molecular pathways that might be involved in determining the extent of response to ivermectins anti-fecundity effect. These data reveal a more complex than previously described pattern of genetic diversity among O. volvulus populations that differ in their geography and response to ivermectin treatment.

genomics

Network based conditional genome wide association analysis of human metabolomics

BackgroundGenome-wide association studies (GWAS) have identified hundreds of loci influencing complex human traits, however, their biological mechanism of action remains mostly unknown. Recent accumulation of functional genomics ( omics) including metabolomics data opens up opportunities to provide a new insight into the functional role of specific changes in the genome. Functional genomic data are characterized by high dimensionality, presence of (strong) statistical dependencies between traits, and, potentially, complex genetic control. Therefore, analysis of such data asks for development of specific statistical genetic methods.\n\nResultsWe propose a network-based, conditional approach to evaluate the impact of genetic variants on omics phenotypes (conditional GWAS, cGWAS). For each trait of interest, based on biological network, we select a set of other traits to be used as covariates in GWAS. The network could be reconstructed either from biological pathway databases or directly from the data. We evaluated our approach using data from a population-based KORA study (n=1,784, 1.7 M SNPs) with measured metabolomics data (151 metabolites) and demonstrated that our approach allows for identification of up to five additional loci not detected by conventional GWAS. We show that this gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.\n\nConclusionsWe demonstrate that in context of metabolomics network-based, conditional genome-wide association analysis is able to dramatically increase power of identification of loci with specific contra-intuitive pleiotropic architecture. Our method has modest computational costs, can utilize summary level GWAS data, and is applicable to other omics data types. We anticipate that application of our method to new and existing data sets will facilitate progress in understanding genetic bases of control of molecular and complex phenotypes.\n\nShort abstractWe propose a network-based, conditional approach for genome-wide analysis of multivariate omics phenotypes. Our methods can incorporate prior biological knowledge about biological pathways from external sources. We evaluated our approach using metabolomics data and demonstrated that our approach has bigger power and allows for identification of additional loci. We show that gain in power is achieved through increased precision of genetic effect estimates, and in presence of specific contra-intuitive pleiotropic scenarios (when genetic and environmental sources of covariance are acting in opposite manner). We justify existence of such scenarios, and discuss possible applications of our method beyond metabolomics.

genomics