bioRxiv ScienceSearch

Biology subjects

Hayes, J.

Publications and source records attributed to Hayes, J..

2 recordsLinked to original sources

Data-adaptive pipeline for filtering and normalizing metabolomics data.

IntroductionUntargeted metabolomics datasets contain large proportions of uninformative features and are affected by a variety of nuisance technical effects that can bias subsequent statistical analyses. Thus, there is a need for versatile and data-adaptive methods for filtering and normalizing data prior to investigating the underlying biological phenomena.\n\nObjectivesHere, we propose and evaluate a data-adaptive pipeline for metabolomics data that are generated by liquid chromatography-mass spectrometry platforms.\n\nMethodsOur data-adaptive pipeline includes novel methods for filtering features based on blank samples, proportions of missing values, and estimated intra-class correlation coefficients. It also incorporates a variant of k-nearest-neighbor imputation of missing values. Finally, we adapted an RNA-Seq approach and R package, scone, to select an appropriate normalization scheme for removing unwanted variation from metabolomics datasets.\n\nResultsUsing two metabolomics datasets that were generated in our laboratory from samples of human blood serum and neonatal blood spots, we compared our data-adaptive pipeline with a traditional filtering and normalization scheme. The data-adaptive approach outperformed the traditional pipeline in almost all metrics related to removal of unwanted variation and maintenance of biologically relevant signatures. The R code for running the data-adaptive pipeline is provided with an example dataset at https://github.com/courtneyschiffman/Data-adaptive-metabolomics.\n\nConclusionOur proposed data-adaptive pipeline is intuitive and effectively reduces technical noise from untargeted metabolomics datasets. It is particularly relevant for interrogation of biological phenomena in data derived from complex matrices associated with biospecimens.

bioinformatics

Validation of Prostate Cancer Risk Variants by CRISPR/Cas9 Mediated Genome Editing

GWAS have identified numerous SNPs associated with prostate cancer risk. One such SNP is rs10993994. It is located in the MSMB promoter, associates with MSMB encoded {beta}-microseminoprotein prostate secretion levels, and is associated with mRNA expression changes in MSMB and the adjacent gene NCOA4. In addition, our previous work showed a second SNP, rs7098889, is in LD with rs10993994 and associated with MSMB expression independent of rs10993994. Here, we generate a series of clones with single alleles removed by double guide RNA (gRNA) mediated CRISPR/Cas9 deletions, through which we demonstrate that each of these SNPs independently and greatly alters MSMB expression in an allele-specific manner. We further show that these SNPs have no substantial effect on the expression of NCOA4. These data demonstrate that a single SNP can have a large effect on gene expression and illustrate the importance of functional validation to deconvolute observed correlations. The method we have developed is generally applicable to test any SNP for which a relevant heterozygous cell line is available.\n\nAuthor summaryIn pursuing the underlying biological mechanism of prostate cancer pathogenesis, scientists utilized the existence of common single nucleotide polymorphisms (SNPs) in human genome as genetic markers to perform large scale genome wide association studies (GWAS) and have so far identified more than a hundred prostate cancer risk variants. Such variants provide an unbiased and systematic new venue to study the disease mechanism, and the next big challenge is to translate these genetic associations to the causal role of altered gene function in oncogenesis. The majority of these variants are waiting to be studied and lots of them may act in oncogenesis through gene expression regulation. To prove the concept, we took rs10993994 and its linked rs7098889 as an example and engineered single cell clones by allelic-specific CRISPR/Cas9 deletion to separate the effect of each allele. We observed that a single nucleotide difference would lead to surprisingly high level of MSMB gene expression change in a gene specific and tissue specific manner. Our study strongly supports the notion that differential level of gene expression caused by risk variants and their associated genetic locus play a major role in oncogenesis and also highlights the importance of studying the function of MSMB encoded {beta}-MSP in prostate cancer pathogenesis.

genetics