bioRxiv ScienceSearch

Biology subjects

Lenz, S.

Publications and source records attributed to Lenz, S..

4 recordsLinked to original sources

Updating genome annotation for the microbial cell factory Aspergillus niger using gene co-expression networks

A significant challenge in our understanding of biological systems is the high number of genes with unknown function in many genomes. The fungal genus Aspergillus contains important pathogens of humans, model organisms, and microbial cell factories. Aspergillus niger is used to produce organic acids, proteins, and is a promising source of new bioactive secondary metabolites. Out of the 14,165 open reading frames predicted in the A. niger genome of only 2% have been experimentally verified and over 6,000 are hypothetical. Here we show that gene co-expression network analysis can be used to overcome this limitation. A meta-analysis of 155 transcriptomics experiments generated co-expression networks for 9,579 genes ([~]65%) of the A. niger genome. By populating this dataset with over 1,200 gene functional experiments from the genus Aspergillus and performing gene ontology enrichment, we could infer biological processes for 9,263 of A. niger genes, including 2,970 hypothetical genes. Experimental validation of selected co-expression sub-networks uncovered four transcription factors involved in secondary metabolite synthesis, which were used to activate production of multiple natural products. This study constitutes a significant step towards systems-level understanding of A. niger, and the datasets can be used to fuel discoveries of model systems, fungal pathogens, and biotechnology.

systems biology

In-Search Selection of Monoisotopic Peaks Improves the Identification of Cross-Linked Peptides

Cross-linking/mass spectrometry (CLMS) has undergone a maturation process akin to standard proteomics by adapting key methods such as false discovery rate control and quantification. A seldom-used search setting in proteomics is the consideration of multiple (lighter) alternative values for the monoisotopic precursor mass to compensate for possible misassignments of the monoisotopic peak. Here, we show that monoisotopic peak assignment is a major weakness of current data handling approaches in cross-linking. Cross-linked peptides often have high precursor masses, which reduces the presence of the monoisotopic peak in the isotope envelope. Paired with generally low peak intensity, this generates a challenge that may not be completely solvable by precursor mass assignment routines. We therefore took an alternative route by in-search assignment of the monoisotopic peak in Xi (Xi-MPA), which considers multiple precursor masses during database search. We compare and evaluate the performance of established preprocessing workflows that partly correct the monoisotopic peak and Xi-MPA on three publicly available datasets. Xi-MPA always delivered the highest number of identifications with ~2 to 4-fold increase of PSMs without compromising identification accuracy as determined by FDR estimation and comparison to crystallographic models.

bioinformatics

A deep learning approach for uncovering lung cancer immunome patterns

Tumor immune cell infiltration is a well known factor related to survival of cancer patients. This has led to deconvolution approaches that can quantify immune cell proportions for each individual. What is missing, is an approach for modeling joint patterns of different immune cell types. We adapt a deep learning approach, deep Boltzmann machines (DBMs), for modeling immune cell gene expression patterns in lung adenocarcinoma. Specifically, a partially partitioned training approach for dealing with a relatively large number of genes. We also propose a sampling-based approach that smooths the original data according to a trained DBM and can be used for visualization and clustering. The identified clusters can subsequently be judged with respect to association with clinical characteristics, such as tumor stage, providing an external criterion for selecting DBM network architecture and tuning parameters for training. We show that the hidden nodes of the trained networks cannot only be linked to clinical characteristics but also to specific genes, which are the visible nodes of the network. We find that hidden nodes that are linked to tumor stage and survival represent expression of T-cell and mast cell genes among others, probably reflecting specific immune cell infiltration patterns. Thus, DBMs, trained and selected by the proposed approach, might provide a useful tool for extracting immune cell gene expression patterns. In the case of lung adenocarcinomas, these patterns are linked to survival as well as other patient characteristics, which could be useful for uncovering the underlying biology.

bioinformatics

Partitioned learning of deep Boltzmann machines for SNP data

Learning the joint distributions of measurements, and in particular identification of an appropriate low-dimensional manifold, has been found to be a powerful ingredient of deep leaning approaches. Yet, such approaches have hardly been applied to single nucleotide polymorphism (SNP) data, probably due to the high number of features typically exceeding the number of studied individuals. After a brief overview of how deep Boltzmann machines (DBMs), a deep learning approach, can be adapted to SNP data in principle, we specifically present a way to alleviate the dimensionality problem by partitioned learning. We propose a sparse regression approach to coarsely screen the joint distribution of SNPs, followed by training several DBMs on SNP partitions that were identified by the screening. Aggregate features representing SNP patterns and the corresponding SNPs are extracted from the DBMs by a combination of statistical tests and sparse regression. In simulated case-control data, we show how this can uncover complex SNP patterns and augment results from univariate approaches, while maintaining type 1 error control. Time-to-event endpoints are considered in an application with acute myeloid lymphoma patients, where SNP patterns are modeled after a pre-screening based on gene expression data. The proposed approach identified three SNPs that seem to jointly influence survival in a validation data set. This indicates the added value of jointly investigating SNPs compared to standard univariate analyses and makes partitioned learning of DBMs an interesting complementary approach when analyzing SNP data.

bioinformatics