bioRxiv ScienceSearch

Biology subjects

Rousu, J.

Publications and source records attributed to Rousu, J..

2 recordsLinked to original sources

Principal Metabolic Flux Mode Analysis

MotivationIn the analysis of metabolism using omics data, two distinct and complementary approaches are frequently used: Principal component analysis (PCA) and Stoichiometric flux analysis. PCA is able to capture the main modes of variability in a set of experiments and does not make many prior assumptions about the data, but does not inherently take into account the flux mode structure of metabolism. Stoichiometric flux analysis methods, such as Flux Balance Analysis (FBA) and Elementary Mode Analysis, on the other hand, produce results that are readily interpretable in terms of metabolic flux modes, however, they are not best suited for exploratory analysis on a large set of samples.\n\nResultsWe propose a new methodology for the analysis of metabolism, called Principal Metabolic Flux Mode Analysis (PMFA), which marries the PCA and Stoichiometric flux analysis approaches in an elegant regularized optimization framework. In short, the method incorporates a variance maximization objective form PCA coupled with a Stoichiometric regularizer, which penalizes projections that are far from any flux modes of the network. For interpretability, we also introduce a sparse variant of PMFA that favours flux modes that contain a small number of reactions. Our experiments demonstrate the versatility and capabilities of our methodology.\n\nAvailabilityMatlab software for PMFA and SPMFA is available in https://github.com/ aalto-ics-kepaco/PMFA.\n\nContactsahely@iitpkd.ac.in, juho.rousu@aalto.fi, Peter.Blomberg@vtt.fi, Sandra.Castillo@vtt.fi\n\nSupplementary informationDetailed results are in Supplementary files. Supplementary data are available at https://github.com/aalto-ics-kepaco/PMFA/blob/master/Results.zip.

bioinformatics

Predicting Protein Producibility In Filamentous Fungi

In this paper we study the problem of predicting the producibility of recombinant proteins in filamentous fungi, especially T. reesei, using machine learning methods. We train supervised and semi-supervised support vector machines with protein sequences, represented by their amino acid composition as well as protein family and domain information. Our results indicate, somewhat surprisingly, that quite modest amount of proteins with experimental data are required to build a state-of-the-art classifier and that additional unlabeled sequences in semi-supervised models do not bring increased predictive performance. Our experiments in cross-species prediction show that models trained for the filamentous fungus A. niger protein dataset can be generalized to predict protein producibility in T. reesei, and vice versa, without sacrificing too much accuracy, regardless of their approximately 500 millions years of divergence. However, predictors trained on E. coli and S. cerevisiae datasets gave variable performance when applied to the filamentous fungi datasets, indicating that while protein producibility prediction can be generalized accross related species, fully generic prediction tools applicable to any protein production host may not be realistic to achieve.

bioinformatics