bioRxiv ScienceSearch

Biology subjects

Sampath, G.

Publications and source records attributed to Sampath, G..

2 recordsLinked to original sources

A minimalist approach to protein identification

Computations on proteome sequence databases show that most proteins can be identified from a proteins isoelectric point (IEP) and digitized linear sequence volume (equal to the total volume of its residues). This is illustrated with four proteomes: H. pylori (1553 proteins), E. coli (4306 proteins), S. cerevisiae (6721 proteins), and H. sapiens (20207 proteins); the identification rate exceeds 90% in all four cases for appropriate parameter values. IEP can be obtained with 1-d gel electrophoresis (GE), whose accuracy is better than 0.01. Linear protein sequence volumes of unbroken proteins can be obtained with a sub-nanometer diameter nanopore that can measure residue volume with a resolution of 0.07-0.1 nm3 (Kennedy et al., Nature Nanotech., 2016, 11, 968-976; Dong et al., ACS Nano, 2017, doi: 10.1021/acsnano.6b08452); the blockade current due to a translocating protein is roughly proportional to the volume it excludes in the pore. There is no need to identify any of the residues. More than 90% of all the proteins have estimated translocation times higher than 1 s, which is within the time resolution of available detectors. This is a minimalist proteolysis-free GE-and nanopore-based single-molecule approach requires very small samples, is non-destructive (the sample can be recovered for reuse), and can be translated with currently available technology into a portable device for possible use in the field, an academic lab, or a pre-screening step preceding conventional mass spectrometry.

bioinformatics

Protein Fingerprinting With A Binary Alphabet

Abstract.If protein sequences are recoded with a binary alphabet derived from a division of the 20 amino acids into two subsets, a protein can be identified from its subsequences by searching through a recoded sequence database. A binary-coded primary sequence can be obtained for an unbroken protein molecule from current blockades in a nanopore. Only two (instead of 20) blockade levels need to be recognized to identify a residues subset; a hard or soft detector can do this with two current thresholds. Computations were done on the complete proteome of Helicobacter pylori (http://www.uniprot.org; database id UP000000210, 1553 sequences) using a binary alphabet based on published data for residue volumes in the range [~]0.06 nm3 to [~]0.225 nm3. Assuming normally distributed volumes, more than 93% of binary subsequences of length 20 from the primary sequences of H. pylori are correct with a confidence level of 90-95%; they can uniquely identify over 98% of the proteins. Recently published work shows that a 0.7 nm diameter nanopore can measure residue volume with a resolution of [~]0.07 nm3; this makes the procedure described here both feasible and practical. This is a non-destructive single-molecule method without the vagaries of proteolysis.

bioengineering