bioRxiv · 10.1101/2020.07.13.201459
Semi-Supervised Learning of Protein Secondary Structure from Single Sequences
Abstract
Accurate modelling of a single orphan protein sequence in the absence of homology information has remained a challenge for several decades. Although not as performant as their homology-based counterparts, single-sequence bioinformatic methods are not constrained by the requirement of evolutionary information and so have a swathe of applications and uses. By taking a bioinformatics approach to semi-supervised machine learning we develop Profile Augmentation of Single Sequences (PASS), a simple but powerful framework for developing accurate single-sequence methods. To demonstrate the effectiveness of PASS we apply it to the mature field of secondary structure prediction. In doing so we develop S4PRED, the successor to the open-source PSIPRED-Single method, which achieves an unprecedented Q3 score of 75.3% on the standard CB513 test. PASS provides a blueprint for the development of a new generation of predictive methods, advancing our ability to model individual protein sequences.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Moffat, L., Jones, D. T.. 2020-07-14. Semi-Supervised Learning of Protein Secondary Structure from Single Sequences. https://doi.org/10.1101/2020.07.13.201459
Cite the original work for its findings. Save a collection to share your selection of sources.