bioRxiv ScienceSearch

Biology subjects

Mamitsuka, H.

Publications and source records attributed to Mamitsuka, H..

3 recordsLinked to original sources

NetGO: Improving Large-scale Protein Function Prediction with Massive Network Information

Automated function prediction (AFP) of proteins is of great significance in biology. In essence, AFP is a large-scale multi-label classification over pairs of proteins and GO terms. Existing AFP approaches, however, have their limitations on both sides of proteins and GO terms. Using various sequence information and the robust learning to rank (LTR) framework, we have developed GOLabeler, a state-of-the-art approach of CAFA3, which overcomes the limitation of the GO term side, such as imbalanced GO terms. Unfortunately, for the protein side issue, available abundant protein information, except for sequences, have not been effectively used for large-scale AFP in CAFA. We propose NetGO that is able to improve large-scale AFP with massive network information. The novelties of NetGO have threefold in using network information: 1) the powerful LTR framework of NetGO efficiently and effectively integrates both sequence and network information, which can easily make large-scale AFP; 2) NetGO can use whole and massive network information of all species (>2000) in STRING (other than only high confidence links and/or some specific species); and 3) NetGO can still use network information to annotate a protein by homology transfer even if it is not covered in STRING. Under numerous experimental settings, we examined the performance of NetGO, such as general performance comparison, species-specific prediction, and prediction on difficult proteins, by using training and test data separated by time-delayed settings of CAFA. Experimental results have clearly demonstrated that NetGO outperforms GOLabeler, DeepGO, and other compared baseline methods significantly. In addition, several interesting findings from our experiments on NetGO would be useful for future AFP research.

bioinformatics

Modelling GxE with historical weather information improves genomic prediction in new environments

Interaction between the genotype and the environment (GxE) has a strong impact on the yield of major crop plants. Although influential, taking GxE explictily into account in plant breeding has remained difficult. Recently GxE has been predicted from environmental and genomic covariates, but existing works have not shown that generalization to new environments and years without access to in-season data is possible and practical applicability remains unclear. Using data from a Barley breeding program in Finland, we construct an in-silico experiment to study the viability of GxE prediction under practical constraints. We show that the response to the environment of a new generation of untested Barley cultivars can be predicted in new locations and years using genomic data, machine learning and historical weather observations for the new locations. Our results highlight the need for models of GxE: non-linear effects clearly dominate linear ones and the interaction between the soil type and daily rain is identified as the main driver for GxE for Barley in Finland. Our study implies that genomic selection can be used to capture the yield potential in GxE effects for future growth seasons, providing a possible means to achieve yield improvements, needed for feeding the growing population.

bioinformatics

GOLabeler: Improving Sequence-based Large-scale Protein Function Prediction by Learning to Rank

Motivation: Gene Ontology (GO) has been widely used to annotate functions of proteins and understand their biological roles. Currently only {inverted exclamation}1% of more than 70 million proteins in UniProtKB have experimental GO annotations, implying the strong necessity of automated function prediction (AFP) of proteins, where AFP is a hard multi-label classification problem due to one protein with a diverse number of GO terms. Most of these proteins have only sequences as input information, indicating the importance of sequence-based AFP (SAFP: sequences are the only input). Furthermore, homology-based SAFP tools are competitive in AFP competitions, while they do not necessarily work well for so-called difficult proteins, which have {inverted exclamation}60% sequence identity to proteins with annotations already. Thus, the vital and challenging problem now is to develop a method for SAFP, particularly for difficult proteins.\n\nMethods: The key of this method is to extract not only homology information but also diverse, deep-rooted information/evidence from sequence inputs and integrate them into a predictor in an efficient and also effective manner. We propose GOLabeler, which integrates five component classifiers, trained from different features, including GO term frequency, sequence alignment, amino acid trigram, domains and motifs, and biophysical properties, etc., in the framework of learning to rank (LTR), a new paradigm of machine learning, especially powerful for multi-label classification.\n\nResults: The empirical results obtained by examining GOLabeler extensively and thoroughly by using large-scale datasets revealed numerous favorable aspects of GOLabeler, including significant performance advantage over state-of-the-art AFP methods.\n\nContact: zhusf@fudan.edu.cn

bioinformatics