bioRxiv ScienceSearch

Biology subjects

Lengauer, T.

Publications and source records attributed to Lengauer, T..

5 recordsLinked to original sources

Relative Principal Components Analysis: Application to Analyzing Biomolecular Conformational Changes

A new method termed \"Relative Principal Components analysis\" (RPCA) is introduced that extracts optimal relevant principal components to describe the change between two data samples representing two macroscopic states. The method is widely applicable in data-driven science. Calculating the components is based on a unified physical framework which introduces the objective function, namely the Kullback-Leibler divergence, appropriate for quantifying the change of the macroscopic state as it is effected by the microscopic features. To demonstrate the applicability of RPCA, we analyze the thermodynamically relevant conformational changes of the protein HIV-1 protease upon binding to different drug molecules. In this case, the RPCA method provides a sound thermodynamic foundation for the analysis of the binding process. The relevant collective (global) conformational changes can be reconstructed from the informative latent variables to exhibit both the enhanced and the restricted conformational fluctuations upon ligand association. Moreover, RPCA characterizes the locally relevant conformational changes which can be presented on the structure of the protein.

biophysics

Epigenome-based prediction of gene expression across species

BackgroundCross-species studies of epigenetic regulation have great potential, yet most epige-nome mapping has focused on human, mouse, and a small number of other model organisms. Here we explore whether existing reference epigenome collections can be leveraged for analyzing other species, by extrapolation and predictive transfer of epigenome information from established model organisms to less well annotated non-model organisms.\n\nResultsWe developed a methodology for cross-species mapping of epigenome data, which we used for predicting tissue-specific gene expression across twelve mammalian and one avian species. Specifically, we trained gradient boosting classifiers to predict gene expression status from reference epigenome data in human and mouse, and we applied these classifiers to epigenome profiles that were computationally transferred between species. The resulting predictions indeed identified tissue-specific differences in gene expression in the target species, thus providing initial validation of the concept of cross-species epigenome extrapolation.\n\nConclusionsOur study establishes a workflow for cross-species epigenome mapping and epigenome-based prediction of gene expression, highlighting the future potential of using epigenome maps from reference species to annotate a potentially large number of target species.

genomics

Integrative analysis of single cell expression data reveals distinct regulatory states in bidirectional promoters

BackgroundBidirectional promoters (BPs) are prevalent in eukaryotic genomes. However, it is poorly understood how the cell integrates different epigenomic information, such as transcription factor (TF) binding and chromatin marks, to drive gene expression at BPs. Single cell sequencing technologies are revolutionizing the field of genome biology. Therefore, this study focuses on the integration of single cell RNA-seq data with bulk ChIP-seq and other epigenetics data, for which single cell technologies are not yet established, in the context of BPs.\n\nResultsWe performed integrative analyses of novel human single cell RNA-seq (scRNA-seq) data with bulk ChIP-seq and other epigenetics data. scRNA-seq data revealed distinct transcription states of BPs that were previously not recognized. We find associations between these transcription states to distinct patterns in structural gene features, DNA accessibility, histone modification, DNA methylation and TF binding profiles.\n\nConclusionsOur results suggest that a complex interplay of all of these elements is required to achieve BP-specific transcriptional output in this specialized promoter configuration. Further, our study implies that novel statistical methods can be developed to deconvolute masked subpopulations of cells measured with different bulk epigenomic assays using scRNA-seq data.

genomics

Partially methylated domains are hallmarks of a cell specific epigenome topology

BackgroundPartially methylated domains, PMDs, are extended regions in the genome exhibiting a reduced average DNA-methylation level. PMDs cover gene-poor and transcriptionally inactive regions and tend to be heterochromatic. Here, we present a first comprehensive comparative analysis of PMDs across more than 190 WGBS methylomes of human and mouse cells providing a deep insight into structural and functional features associated with PMDs.\n\nResultsPMDs are ubiquitous signatures covering up to 75% of the genome in human and mouse cells irrespective of their tissue or cell origin. Additionally, each cell type comes with a distinct set of specific PMDs, and genes expressed in such PMDs show a strong cell type effect. Demethylation strength varies in PMDs with a tendency towards a more pronounced effect in differentiating and replicating cells. The strongest demethylation is observed in highly proliferating and immortal cancer cell lines. A decrease of DNA-methylation within PMDs tends to be linked to an increase in heterochromatic histone marks and a decrease of gene expressions. Characteristic combinations of heterochromatic signatures in PMDs are linked to domains of early, middle and late DNA-replication.\n\nConclusionPMDs are prominent signatures of long-range epigenomic organization. Integrative analysis identifies PMDs as important general, lineage- and cell-type specific topological features. PMD changes are hallmarks of cell differentiation. Demethylation of PMDs combined with increased heterochromatic marks is a feature linked to enhanced cell proliferation. In combination with broad histone marks PMDs demarcate distinct domains of late DNA-replication.

bioinformatics

Combining transcription factor binding affinities with open-chromatin data for accurate gene expression prediction

The binding and contribution of transcription factors (TF) to cell specific gene expression is often deduced from open-chromatin measurements to avoid costly TF ChIP-seq assays. Thus, it is important to develop computational methods for accurate TF binding prediction in open-chromatin regions (OCRs). Here, we report a novel segmentation-based method, TEPIC, to predict TF binding by combining sets of OCRs with position weight matrices. TEPIC can be applied to various open-chromatin data, e.g. DNaseI-seq and NOMe-seq. Additionally, Histone-Marks (HMs) can be used to identify candidate TF binding sites. TEPIC computes TF affinities and uses open-chromatin/HM signal intensity as quantitative measures of TF binding strength. Using machine learning, we find low affinity binding sites to improve our ability to explain gene expression variability compared to the standard presence/absence classification of binding sites. Further, we show that both footprints and peaks capture essential TF binding events and lead to a good prediction performance. In our application, gene-based scores computed by TEPIC with one open-chromatin assay nearly reach the quality of several TF ChIP-seq datasets. Finally, these scores correctly predict known transcriptional regulators as illustrated by the application to novel DNaseI-seq and NOMe-seq data for primary human hepatocytes and CD4+ T-cells, respectively.

bioinformatics