bioRxiv ScienceSearch

Biology subjects

Lee, S.-I.

Publications and source records attributed to Lee, S.-I..

8 recordsLinked to original sources

Explainable machine learning prediction of synergistic drug combinations for precision cancer medicine

Although combination therapy has been a mainstay of cancer treatment for decades, it remains challenging, both to identify novel effective combinations of drugs and to determine the optimal combination for a particular patients tumor. While there have been several recent efforts to test drug combinations in vitro, examining the immense space of possible combinations is far from being feasible. Thus, it is crucial to develop datadriven techniques to computationally identify the optimal drug combination for a patient. We introduce TreeCombo, an extreme gradient boosted tree-based approach to predict synergy of novel drug combinations, using chemical and physical properties of drugs and gene expression levels of cell lines as features. We find that TreeCombo significantly outperforms three other state-of-theart approaches, including the recently developed DeepSynergy, which uses the same set of features to predict synergy using deep neural networks. Moreover, we found that the predictions from our approach were interpretable, with genes having well-established links to cancer serving as important features for prediction of drug synergy.

cancer biology

MD-AD: Multi-task deep learning for Alzheimer’s disease neuropathology

Systematic modeling of Alzheimers Disease (AD) neuropathology based on brain gene expression would provide valuable insights into the disease. However, relative scarcity and regional heterogeneity of brain gene expression and neuropathology datasets obscure the ability to robustly identify expression markers. We propose MD-AD (Multi-task Deep learning for AD) to effectively combine heterogeneous AD datasets by simultaneously modeling multiple phenotypes with shared layers. MD-AD leads to an 8% and 5% reduction in mean squared error over MLP for predicting counts of two AD hallmarks: plaques and tangles. It also leads to a 40% and 30% reduction in classification error over MLP for two common staging systems for AD: CERAD score and Braak stage. Additionally, MD-ADs network representation tends to better capture known metabolic pathways, including some AD-related pathways. Together, these results indicate that MD-AD is particularly useful for learning expressive network representations from heterogeneous and sparsely labeled AD data.

systems biology

A computational framework identifying concordant gene expression-neuropathology associations reveals Complex I as a potential Alzheimer’s disease therapeutic target

Identifying gene expression markers for Alzheimers disease (AD) neuropathology through meta-analysis is a complex undertaking because available data are often from different studies and/or brain regions involving study-specific confounders and/or region-specific biological processes. Here we introduce a novel probabilistic model-based framework, DECODER, leveraging these discrepancies to identify robust biomarkers for complex phenotypes. Our experiments present: (1) DECODERs potential as a general meta-analysis framework widely applicable to various diseases (e.g., AD and cancer) and phenotypes (e.g., Amyloid-{beta} (A{beta}) pathology, tau pathology, and survival), (2) our results from a meta-analysis using 1,746 human brain tissue samples from nine brain regions in three studies -- the largest expression meta-analysis for AD, to our knowledge --, and (3) in vivo validation of identified modifiers of A{beta} toxicity in a transgenic Caenorhabditis elegans model expressing AD-associated A{beta}, which pinpoints mitochondrial Complex I as a critical mediator of proteostasis and a promising pharmacological avenue toward treating AD.

systems biology

Identifying progressive gene network perturbation from single-cell RNA-seq data

Identifying the gene regulatory networks that control development or disease is one of the most important problems in biology. Here, we introduce a computational approach, called PIPER (ProgressIve network PERturbation), to identify the perturbed genes that drive differences in the gene regulatory network across different points in a biological progression. PIPER employs algorithms tailor-made for single cell RNA sequencing (scRNA-seq) data to jointly identify gene networks for multiple progressive conditions. It then performs differential network analysis along the identified gene networks to identify master regulators. We demonstrate that PIPER outperforms state-of-the-art alternative methods on simulated data and is able to predict known key regulators of differentiation on real scRNA-Seq datasets.

bioinformatics

DeepProfile: Deep learning of patient molecular profiles for precision medicine in acute myeloid leukemia

We present the DeepProfile framework, which learns a variational autoencoder (VAE) network from thousands of publicly available gene expression samples and uses this network to encode a low-dimensional representation (LDR) to predict complex disease phenotypes. To our knowledge, DeepProfile is the first attempt to use deep learning to extract a feature representation from a vast quantity of unlabeled (i.e, lacking phenotype information) expression samples that are not incorporated into the prediction problem. We use Deep-Profile to predict acute myeloid leukemia patients in vitro responses to 160 chemotherapy drugs. We show that, when compared to the original features (i.e., expression levels) and LDRs from two commonly used dimensionality reduction methods, DeepProfile: (1) better predicts complex phenotypes, (2) better captures known functional gene groups, and (3) better reconstructs the input data. We show that DeepProfile is generalizable to other diseases and phenotypes by using it to predict ovarian cancer patients tumor invasion patterns and breast cancer patients disease subtypes.

bioinformatics

AIControl: Replacing matched control experiments with machine learning improves ChIP-seq peak identification

ChIP-seq is a technique to determine binding locations of transcription factors, which remains a central challenge in molecular biology. Current practice is to use a \"control\" dataset to remove background signals from a immunoprecipitation (IP) target dataset. We introduce the AlControl framework, which eliminates the need to obtain a control dataset and instead identifies binding peaks by estimating the distributions of background signals from many publicly available control ChIP-seq datasets. We thereby avoid the cost of running control experiments while simultaneously increasing the accuracy of binding location identification. Specifically, AIControl can (1) estimate background signals at fine resolution, (2) systematically weigh the most appropriate control datasets in a data-driven way, (3) capture sources of potential biases that may be missed by one control dataset, and (4) remove the need for costly and time-consuming control experiments. We applied AIControl to 410 IP datasets in the ENCODE ChIP-seq database, using 440 control datasets from 107 cell types to impute background signal. Without using matched control datasets, AIControl identified peaks that were more enriched for putative binding sites than those identified by other popular peak callers that used a matched control dataset. We also demonstrated that our framework identifies binding sites that recover documented protein interactions more accurately.

bioinformatics

Explainable machine learning predictions to help anesthesiologists prevent hypoxemia during surgery

Hypoxemia causes serious patient harm, and while anesthesiologists strive to avoid hypoxemia during surgery, anesthesiologists are not reliably able to predict which patients will have intraoperative hypoxemia. Using minute by minute EMR data from fifty thousand surgeries we developed and tested a machine learning based system called Prescience that predicts real-time hypoxemia risk and presents an explanation of factors contributing to that risk during general anesthesia. Prescience improved anesthesiologists performance when providing interpretable hypoxemia risks with contributing factors. The results suggest that if anesthesiologists currently anticipate 15% of events, then with Prescience assistance they could anticipate 30% of events or an estimated additional 2.4 million annually in the US, a large portion of which may be preventable because they are attributable to modifiable factors. The prediction explanations are broadly consistent with the literature and anesthesiologists prior knowledge. Prescience can also improve clinical understanding of hypoxemia risk during anesthesia by providing general insights into the exact changes in risk induced by certain patient or procedure characteristics. Making predictions of complex medical machine learning models (such as Prescience) interpretable has broad applicability to other data-driven prediction tasks in medicine.

bioinformatics

DeepATAC: A deep-learning method to predict regulatory factor binding activity from ATAC-seq signals

Determining the binding locations of regulatory factors, such as transcription factors and histone modifications, is essential to both basic biology research and many clinical applications. Obtaining such genome-wide location maps directly is often invasive and resource-intensive, so it is common to impute binding locations from DNA sequence or measures of chromatin accessibility. We introduce DeepATAC, a deep-learning approach for imputing binding locations that uses both DNA sequence and chromatin accessibility as measured by ATAC-seq. DeepATAC significantly outperforms current approaches such as FIMO motif predictions overlapped with ATAC-seq peaks, and models based only on DNA sequence, such as DeepSEA. Visualizing the input importances for the DeepATAC model reveals DNA sequence motifs and ATAC-seq signal patterns that are important for predicting binding events. The Keras implementation and analysis pipelines of DeepATAC are available at https://github.com/hiranumn/deepatac.

bioinformatics