bioRxiv ScienceSearch

Biology subjects

Heinig, M.

Publications and source records attributed to Heinig, M..

2 recordsLinked to original sources

MetaMap: An atlas of metatranscriptomic reads in human disease-related RNA-seq data

BackgroundWith the advent of the age of big data in bioinformatics, large volumes of data and high performance computing power enable researchers to perform re-analyses of publicly available datasets at an unprecedented scale. Ever more studies imply the microbiome in both normal human physiology and a wide range of diseases. RNA sequencing technology (RNA-seq) is commonly used to infer global eukaryotic gene expression patterns under defined conditions, including human disease-related contexts, but its generic nature also enables the detection of microbial and viral transcripts.\n\nFindingsWe developed a bioinformatic pipeline to screen existing human RNA-seq datasets for the presence of microbial and viral reads by re-inspecting the non-human-mapping read fraction. We validated this approach by recapitulating outcomes from 6 independent controlled infection experiments of cell line models and comparison with an alternative metatranscriptomic mapping strategy. We then applied the pipeline to close to 150 terabytes of publicly available raw RNA-seq data from >17,000 samples from >400 studies relevant to human disease using state-of-the-art high performance computing systems. The resulting data of this large-scale re-analysis are made available in the presented MetaMap resource.\n\nConclusionsOur results demonstrate that common human RNA-seq data, including those archived in public repositories, might contain valuable information to correlate microbial and viral detection patterns with diverse diseases. The presented MetaMap database thus provides a rich resource for hypothesis generation towards the role of the microbiome in human disease.

bioinformatics

Efficient parameterization of large-scale mechanistic models enables drug response prediction for cancer cell lines

The response of cancer cells to drugs is determined by various factors, including the cells mutations and gene expression levels. These factors can be assessed using next-generation sequencing. Their integration with vast prior knowledge on signaling pathways is, however, limited by the availability of mathematical models and scalable computational methods. Here, we present a computational framework for the parameterization of large-scale mechanistic models and its application to the prediction of drug response of cancer cell lines from exome and transcriptome sequencing data. With this framework, we parameterized a mechanistic model describing major cancer-associated signaling pathways (>1200 species and >2600 reactions) using drug response data. For the parameterized mechanistic model, we found a prediction accuracy, which exceeds that of the considered statistical approaches. Our results demonstrate for the first time the massive integration of heterogeneous datasets using large-scale mechanistic models, and how these models facilitate individualized predictions of drug response. We anticipate our parameterized model to be a starting point for the development of more comprehensive, curated models of signaling pathways, accounting for additional pathways and drugs.

systems biology