bioRxiv ScienceSearch

Biology subjects

Way, G. P.

Publications and source records attributed to Way, G. P..

6 recordsLinked to original sources

Extracting a Biologically Relevant Latent Space from Cancer Transcriptomes with Variational Autoencoders

The Cancer Genome Atlas (TCGA) has profiled over 10,000 tumors across 33 different cancer-types for many genomic features, including gene expression levels. Gene expression measurements capture substantial information about the state of each tumor. Certain classes of deep neural network models are capable of learning a meaningful latent space. Such a latent space could be used to explore and generate hypothetical gene expression profiles under various types of molecular and genetic perturbation. For example, one might wish to use such a model to predict a tumors response to specific therapies or to characterize complex gene expression activations existing in differential proportions in different tumors. Variational autoencoders (VAEs) are a deep neural network approach capable of generating meaningful latent spaces for image and text data. In this work, we sought to determine the extent to which a VAE can be trained to model cancer gene expression, and whether or not such a VAE would capture biologically-relevant features. In the following report, we introduce a VAE trained on TCGA pan-cancer RNA-seq data, identify specific patterns in the VAE encoded features, and discuss potential merits of the approach. We name our method \"Tybalt\" after an instigative, cat-like character who sets a cascading chain of events in motion in Shakespeares \"Romeo and Juliet\". From a systems biology perspective, Tybalt could one day aid in cancer stratification or predict specific activated expression patterns that would result from genetic changes or treatment effects.

bioinformatics

Functional network community detection can disaggregate and filter multiple underlying pathways in enrichment analyses

Differential expression experiments or other analyses often end in a list of genes. Pathway enrichment analysis is one method to discern important biological signals and patterns from noisy expression data. However, pathway enrichment analysis may perform suboptimally in situations where there are multiple implicated pathways - such as in the case of genes that define subtypes of complex diseases. Our simulation study shows that in this setting, standard overrepresentation analysis identifies many false positive pathways along with the true positives. These false positives hamper investigators attempts to glean biological insights from enrichment analysis. We develop and evaluate an approach that combines community detection over functional networks with pathway enrichment to reduce false positives. Our simulation study demonstrates that a large reduction in false positives can be obtained with a small decrease in power. Though we hypothesized that multiple communities might underlie previously described subtypes of high-grade serous ovarian cancer and applied this approach, our results do not support this hypothesis. In summary, applying community detection before enrichment analysis may ease interpretation for complex gene sets that represent multiple distinct pathways.

bioinformatics

PathCORE: Visualizing globally co-occurring pathways in large transcriptomic compendia

BackgroundInvestigators often interpret genome-wide data by analyzing the expression levels of genes within pathways. While this within-pathway analysis is routine, the products of any one pathway can affect the activity of other pathways. Past efforts to identify relationships between biological processes have evaluated overlap in knowledge bases or evaluated changes that occur after specific treatments. Individual experiments can highlight condition-specific pathway-pathway relationships; however, constructing a complete network of such relationships across many conditions requires analyzing results from many studies. ResultsWe developed PathCORE-T framework by implementing existing methods to identify pathway-pathway transcriptional relationships evident across a broad data compendium. PathCORE-T is applied to the output of feature construction algorithms; it identifies pairs of pathways observed in features more than expected by chance as functionally co-occurring. We demonstrate PathCORE-T by analyzing an existing eADAGE model of a microbial compendium and building and analyzing NMF features from the TCGA dataset of 33 cancer types. The PathCORE-T framework includes a demonstration web interface, with source code, that users can launch to (1) visualize the network and (2) review the expression levels of associated genes in the original data. PathCORE-T creates and displays the network of globally co-occurring pathways based on features observed in a machine learning analysis of gene expression data. ConclusionsThe PathCORE-T framework identifies transcriptionally co-occurring pathways from the results of unsupervised analysis of gene expression data and visualizes the relationships between pathways as a network. PathCORE-T recapitulated previously described pathway-pathway relationships and suggested experimentally testable additional hypotheses that remain to be explored.

bioinformatics

Opportunities And Obstacles For Deep Learning In Biology And Medicine

Deep learning, which describes a class of machine learning algorithms, has recently showed impressive results across a variety of domains. Biology and medicine are data rich, but the data are complex and often ill-understood. Problems of this nature may be particularly well-suited to deep learning techniques. We examine applications of deep learning to a variety of biomedical problems--patient classification, fundamental biological processes, and treatment of patients--and discuss whether deep learning will transform these tasks or if the biomedical sphere poses unique challenges. We find that deep learning has yet to revolutionize or definitively resolve any of these problems, but promising advances have been made on the prior state of the art. Even when improvement over a previous baseline has been modest, we have seen signs that deep learning methods may speed or aid human investigation. More work is needed to address concerns related to interpretability and how to best model each problem. Furthermore, the limited amount of labeled data for training presents problems in some domains, as do legal and privacy constraints on work with sensitive health records. Nonetheless, we foresee deep learning powering changes at both bench and bedside with the potential to transform several areas of biology and medicine.

bioinformatics

Reference-free deconvolution of DNA methylation signatures identifies common differentially methylated gene regions on 1p36 across breast cancer subtypes

Breast cancer is a complex disease and studying DNA methylation (DNAm) in tumors is complicated by disease heterogeneity. We compared DNAm in breast tumors with normal-adjacent breast samples from The Cancer Genome Atlas (TCGA). We constructed models stratified by tumor stage and PAM50 molecular subtype and performed cell-type reference-free deconvolution on each model. We identified nineteen differentially methylated gene regions (DMGRs) in early stage tumors across eleven genes (AGRN, C1orf170, FAM41C, FLJ39609, HES4, ISG15, KLHL17, NOC2L, PLEKHN1, SAMD11, WASH5P). These regions were consistently differentially methylated in every subtype and all implicated genes are localized on chromosome 1p36.3. We also validated seventeen DMGRs in an independent data set. Identification and validation of shared DNAm alterations across tumor subtypes in early stage tumors advances our understanding of common biology underlying breast carcinogenesis and may contribute to biomarker development. We also provide evidence on the importance and potential function of 1p36 in cancer.

genomics

Determining causal genes from GWAS signals using topologically associating domains

Genome wide association studies (GWAS) have contributed significantly to the understanding of complex disease genetics. However, GWAS only report associated signals and do not necessarily identify culprit genes. As most signals occur in non-coding regions of the genome, it is often challenging to assign genomic variants to the underlying causal mechanism(s). Topologically associating domains (TADs) are primarily cell-type independent genomic regions that define interactome boundaries and can aid in the designation of limits within which an association most likely impacts gene function. We describe and validate a computational method that uses the genic content of TADs to discover candidate genes. Our method, called \"TAD_Pathways,\" performs a Gene Ontology (GO) analysis over genes that reside within TAD boundaries corresponding to GWAS signals for a given trait or disease. We applied our pipeline to the GWAS catalog entries associated with bone mineral density (BMD), identifying Skeletal System Development (Benjamini-Hochberg adjusted p=1.02x10-5) as the top ranked pathway. In many cases, our method implicated a gene other than the nearest gene. Our molecular experiments describe a novel example: ACP2, implicated at the canonical ARHGAP1 locus. We found ACP2 to be an important regulator of osteoblast metabolism, whereas ARHGAP1 was not supported. Our results via the example of BMD demonstrate how basic principles of three-dimensional genome organization can define biologically informed association windows.

bioinformatics