bioRxiv ScienceSearch

Biology subjects

G Sampath

Publications and source records attributed to G Sampath.

6 recordsLinked to original sources

Peptide partitions and protein identification: a computational analysis

Peptide sequences from a proteome can be partitioned into N mutually exclusive sets and used to identify their parent proteins in a sequence database. This is illustrated with the human proteome (http://www.uniprot.org; id UP000005640), which is partitioned into eight subsets KZ*R, KZ*D, KZ*E, KZ*, Z*R, Z*D, Z*E, and Z*, where Z [isin] {A, N, C, Q, G, H, I, L, M, F, P, S, T, W, Y, V} and Z* {equiv} 0 or more occurrences of Z. If the full peptide sequence is known then over 98% of the proteins in the proteome can be identified from such sequences. The rate exceeds 78% if the positions of four internal residue types are known. When the standard set of 20 amino acids is replaced with an alphabet of size four based on residue volume the identification rate exceeds 96%. In an information-theoretic sense this last result suggests that protein sequences effectively carry nearly the same amount of information as the exon sequences in the genome that code for them using an alphabet of size four. An appendix discusses possible in vitro methods to create peptide partitions and potential ways to sequence partitioned peptides.

Bioinformatics

A digital approach to protein identification and quantity estimation using tandem nanopores, peptidases, and database search

A digital approach to protein identification and quantity estimation using electrical measurements and database search is proposed. It is based on an electrolytic cell with two (three) nanopores and one (two) peptidase(s) covalently attached to the trans side of a pore. An unknown protein is digested by a reagent or peptidase into peptides ending in a known amino acid; the peptides enter the cell, pass through the first pore, and are fragmented by a high-specificity endopeptidase. The second enzyme, if present, is an exopeptidase that cleaves the fragments into single residues after the second pore. Level transitions in an ionic blockade or transverse current pulse due to residues in a fragment or individual pulses due to single residues are counted. This yields the positions of the endopeptidases target in the peptide, and, together with the peptides terminal residue, a partial sequence. Search through the Uniprot database for such sequences identifies over 90% of the proteins in the human proteome. The percentage can be increased by repeating the procedure with other reagents and cells specific to other residues, close to 100% may be possible. Sample purification to homogeneity is not required as the method applies to an arbitrary mixture of proteins; the quantity of a protein in the sample is estimated from the number of identifying peptides sensed over a long run. A Fokker-Planck model gives minimum enzyme turnover intervals required for ordered sensing of peptide fragments. With thick (80-100 nm) pores, required pulse resolution times are within the capability of CMOS detectors. The method can be implemented with existing technology; several related issues are discussed.

Bioengineering

A circuit theory of protein structure

Protein secondary and tertiary structure is modeled as a linear passive analog lumped electrical circuit. Modeling is based on the structural similarity between helix, sheet, turn/loop, and helix pair in proteins and inductor, capacitor, resistor, and transformer in electrical circuits; it includes methods from circuit analysis and synthesis. A protein circuit is a one-port with a restrictive circuit topology (for example, the circuit for a secondary structure cannot be a Foster II ladder or a Wheatstone-like bridge). It has a rational positive real impedance function whose pole-zero distribution serves as a compact descriptor of secondary and tertiary structure, which is reminiscent of the Ramachandran plot. Standard circuit analysis methods such as node/loop equations and pole-zero maps may be used to study differences at the secondary and tertiary levels within and across proteins. Pairs of interacting proteins can be modeled as two-ports and studied via transfer functions. Similarly circuit synthesis methods can be used to construct protein circuits whose real counterparts may or may not exist. An analysis example shows how a protein circuit is constructed for thioredoxin and its pole-zero map obtained. A synthesis example shows how an electrical circuit with a single Brune section is obtained from a specified set of poles and zeros and then mapped to an artificial protein with a helix pair (corresponding to the transformer in the Brune section). Possible applications to folding, drug design, and visualization are indicated.

Bioengineering

A quasi-digital approach to peptide sequencing using tandem nanopores with endo- and exo-peptidases

A method of sequencing peptides using tandem cells (RSC Adv., 2015, 5, 167-171; RSC Adv., 2015, 5, 30694-30700) and peptidases is considered. A double tandem cell (two tandem cells in tandem) has three nanopores in series, an amino-acid-specific endopeptidase attached downstream of the first pore, and an exopeptidase attached downstream of the second pore. The endopeptidase cleaves a peptide threaded through the first pore into fragments that are well separated in time. Fragments pass through the second pore and are each cleaved by the exopeptidase into a series of single residues; the latter pass through the third pore and cause separate current blockades that can be counted. This leads to an ordered list of integers corresponding to the number of residues in each fragment. With 20 cells, one per amino acid type, and 20 peptide copies, the resulting 20 lists of integers are used by a simple algorithm to assemble the sequence. This is a quasi-digital process that uses the lengths of subsequences to sequence the peptide, it differs from conventional analog methods which seek to identify monomers in a polymer via differences in blockade levels, residence times, or transverse currents. Several implementation issues are discussed. In particular the problem of fast analyte translocation, widely considered intransigent, may be resolved through the use of a sufficiently long (40-60 nm) third pore. This translates to a required bandwidth of 1-2 MHz, which is within the range of currently available CMOS circuits.

Bioengineering

Peptide sequencing in an electrolytic cell with two nanopores in tandem and exopeptidase

A nanopore-based approach to peptide sequencing without labels or immobilization is considered. It is based on a tandem cell (RSC Adv., 2015, 5, 167-171) with the structure [cis1, upstream pore (UNP), trans1/cis2, downstream pore (DNP), trans2]. An amino or carboxyl exopeptidase attached to the downstream side of UNP cleaves successive leading residues in a peptide threading from cis1 through UNP. A cleaved residue translocates to and through DNP where it is identified. A Fokker-Planck model is used to compute translocation statistics for each amino acid type. Multiple discriminators, including a variant of the current blockade level and translocation times through trans1/cis2 and DNP, identify a residue. Calculations show the 20 amino acids to be grouped by charge (+, -, neutral) and ordered within each group (which makes error correction easier). The minimum cleaving interval required of the exopeptidase, the sample size (number of copies of the peptide to sequence or runs with one copy) to identify a residue with a given confidence level, and confidence levels for a given sample size are calculated. The results suggest that if the exopeptidase cleaves each and every residue and does so in a reasonable time, peptide sequencing with acceptable (and correctable) errors may be feasible. If validated experimentally the proposed device could be an alternative to mass spectrometry and gel electrophoresis. Implementation-related issues are discussed.

Bioengineering

A Tandem Cell for Nanopore-based DNA Sequencing with Exonuclease

A tandem cell is proposed for DNA sequencing in which an exonuclease enzyme cleaves bases (mononucleotides) from a strand of DNA for identification inside a nanopore. It has two nanopores and three compartments with the structure [cis1, upstream nanopore (UNP), trans1 = cis2, downstream nanopore (DNP), trans2]. The exonuclease is attached to the exit side of UNP in trans1/cis2. A cleaved base cannot regress into cis1 because of the remaining DNA strand in UNP. A profiled electric field over DNP with positive and negative components slows down base translocation through DNP. The proposed structure is modeled with a Fokker-Planck equation and a piecewise solution presented. Results from the model indicate that with probability approaching 1 bases enter DNP in their natural order, are detected without any loss, and do not regress into DNP after progressing into trans2. Sequencing efficiency with a tandem cell would then be determined solely by the level of discrimination among the base types inside DNP.

Bioengineering