bioRxiv ScienceSearch

Biology subjects

Nadav Brandes

Publications and source records attributed to Nadav Brandes.

2 recordsLinked to original sources

ASAP: A Machine-Learning Framework for Local Protein Properties

Determining residue level protein properties, such as the sites for post-translational modifications (PTMs) are vital to understanding proteins at all levels of function. Experimental methods are costly and time-consuming, thus high confidence predictions become essential for functional knowledge at a genomic scale. Traditional computational methods based on strict rules (e.g. regular expressions) fail to annotate sites that lack substantial similarity. Thus, Machine Learning (ML) methods become fundamental in annotating proteins with unknown function. We present ASAP (Amino-acid Sequence Annotation Prediction), a universal ML framework for residue-level predictions. ASAP extracts efficiently and fast large set of window-based features from raw sequences. The platform also supports easy integration of external features such as secondary structure or PSSM profiles. The features are then combined to train underlying ML classifiers. We present a detailed case study for ASAP that was used to train CleavePred, a state-of-the-art protein precursor cleavage sites predictor. Protein cleavage is a fundamental PTM shared by a wide variety of protein groups with minimal sequence similarity. Current computational methods have high false positive rates, making them suboptimal for this task. CleavePred has a simple Python API, and is freely accessible via a web-based application. The high performance of ASAP toward the task of precursor cleavage is suited for analyzing new proteomes at a genomic scale. The tool is attractive to protein design, mass spectrometry search engines and the discovery of new peptide hormones. In summary, we illustrate ASAP as an entry point for predicting PTMs. The approach and flexibility of the platform can easily be extended for additional residue specific tasks. ASAP and CleavePred source code available at https://github.com/ddofer/asap.

Bioinformatics

Overlapping Genes and Size Constraints in Viruses - An Evolutionary Perspective

Viruses are the simplest replicating units, characterized by a limited number of coding genes and an exceptionally high rate of overlapping genes. We sought a unified explanation for the evolutionary constraints that govern genome sizes, gene overlapping and capsid properties. We performed an unbiased statistical analysis over the [~]100 known viral families, and came to refute widespread assumptions regarding viral evolution. We found that the volume utilization of viral capsids is often low, and greatly varies among families. Most notably, we show that the total amount of gene overlapping is tightly bounded. Although viruses expand three orders of magnitude in genome length, their absolute amount of gene overlapping almost never exceeds 1500 nucleotides, and mostly confined to <4 significant overlapping instances. Our results argue against the common theory by which gene overlapping is driven by a necessity of viruses to compress their genome. Instead, we support the notion that overlapping has a role in gene novelty and evolution exploration.

Evolutionary Biology