bioRxiv ScienceSearch

Biology subjects

Samudrala, R.

Publications and source records attributed to Samudrala, R..

4 recordsLinked to original sources

CANDOCK: Chemical atomic network based hierarchical flexible docking algorithm using generalized statistical potentials

Small molecule docking has proven to be invaluable for drug design and discovery. However, existing docking methods have several limitations, such as, improper treatment of the interactions of essential components in the chemical environment of the binding pocket (e.g. cofactors, metal-ions, etc.), incomplete sampling of chemically relevant ligand conformational space, and the inability to consistently correlate docking scores of the best binding pose with experimental binding affinities. We present CANDOCK, a novel docking algorithm that utilizes a hierarchical approach to reconstruct ligands from an atomic grid using graph theory and generalized statistical potential functions to sample biologically relevant ligand conformations. Our algorithm accounts for protein flexibility, solvent, metal ions and cofactors interactions in the binding pocket that are traditionally ignored by current methods. We evaluate the algorithm on the PDBbind and Astex proteins to show its ability to reproduce the binding mode of the ligands that is independent of the initial ligand conformation in these benchmarks. Finally, we identify the best selector and ranker potential functions, such that, the statistical score of best selected docked pose correlates with the experimental binding affinities of the ligands for any given protein target. Our results indicate that CANDOCK is a generalized flexible docking method that addresses several limitations of current docking methods by considering all interactions in the chemical environment of a binding pocket for correlating the best docked pose with biological activity.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=28 SRC=\"FIGDIR/small/442897v2_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (16K):\norg.highwire.dtl.DTLVardef@1c1b2b0org.highwire.dtl.DTLVardef@1ba66a4org.highwire.dtl.DTLVardef@a8bfd3org.highwire.dtl.DTLVardef@c9d3bb_HPS_FORMAT_FIGEXP M_FIG C_FIG

biophysics

Identifying protein subsets and features responsible for improved drug repurposing accuracies using the CANDO platform

Drug repurposing is a valuable tool for combating the slowing rates of novel therapeutic discovery. The Computational Analysis of Novel Drug Opportunities (CANDO) platform performs shotgun repurposing of 3,733 drugs/compounds that map to 2,030 indications/diseases by predicting their interactions with 46,784 protein structures and relating them via proteomic interaction signatures. The accuracy of the CANDO platform is evaluated using our benchmarking protocol that assesses indication accuracies based on whether or not pairs of drugs associated with the same indication can be captured within a certain cutoff, which is a measure of the drug repurposing recovery rate. To identify subsets of proteins that exhibit the same therapeutic effectiveness as the full set, groups of 8 proteins were randomly selected and subsequently benchmarked 50 times. The resulting protein sets were ranked according to average indication accuracy, pairwise accuracy, and coverage (count of indications with non-zero accuracy). The best 50 subsets of 8 according to each metric were progressively combined into supersets after each iteration and benchmarked. These supersets yield up to 14% improvement in benchmarking accuracy, and represent a 100-1,000 fold reduction in the number of proteins relative to the full set. Protein supersets optimized using independent compound libraries derived from the full library were cross-tested and were shown to reproduce the performance relative to using all 46,784 proteins, indicating that these reduced size supersets are broadly applicable for characterizing drug behavior. Further analysis revealed that sets comprised of proteins with more equitably diverse ligand interactions are important for describing drug behavior. Our work elucidates the role of particular protein subsets and corresponding ligand interactions that play a role in computational drug repurposing, and paves the way for the use of machine learning approaches to further improve the accuracy of the CANDO platform and its repurposing potential.\n\nAuthor summaryDrug repurposing is a valuable approach for ameliorating the current problems plaguing drug discovery. We introduce a novel protein subset analysis pipeline that allows us to elucidate features important for drug repurposing accuracies using the Computational Analysis of Novel Drug Opportunities (CANDO) platform. Our platform relates drugs based on the similarity of their interactions with a diverse library of proteins. We subjected all proteins in the platform to a splitting and ranking protocol that ranked protein subsets based on their benchmarking performance. Further analysis of the best performing protein subsets revealed that the most useful proteins for describing how small molecule compounds behave in biological systems are those that are predicted to interact with a structurally diverse range of ligands. We hypothesize that this is a consequence of the multitarget nature of drugs and, conversely, the implied promiscuity of proteins in biological systems. These results may be used to make drug discovery more accurate and efficient by alleviating some of its bottlenecks, bringing us one step further in better understanding how drugs behave in the context of their environments.

bioinformatics

Accurate informatic modeling of tooth enamel pellicle interactions by training substitution matrices with Mat4Pep

MotivationProtein-hydroxyapatite interactions govern the development and homeostasis of teeth and bone. Characterization would enable design of peptides to regenerate mineralized tissues and control attachments such as ligaments and dental plaque. Progress has been limited because no available methods produce robust data for assessing phase interfaces.\n\nResultsWe show that tooth enamel pellicle peptides contain subtle sequence similarities that encode hydroxyapatite binding mechanisms, by segregating pellicle peptides from control sequences using our previously developed substitution matrix-based peptide comparison protocol (Oren et al., 2007), with improvements. Sampling diverse matrices, adding biological control sequences, and optimizing matrix refinement algorithms improves discrimination from 0.81 to 0.99 AUC in leave-one-out experiments. Other contemporary methods fail on this problem. We find hydroxyapatite interaction sequence patterns by applying the resulting selected refined matrix (\"pellitrix\") to cluster the peptides and build subgroup alignments. We identify putative hydroxyapatite maturation domains by application to enamel biomineralization proteins and prioritize putative novel pellicle peptides identified by In stageTip (iST) mass spectrometry. The sequence comparison protocol outperforms other contemporary options for this small and heterogeneous group, and is generalized for application to any group of peptides.\n\nAvailabilitySoftware to apply this protocol is freely available at github.com/JeremyHorst/Mat4Pep and compbio.org/protinfo/ Mat4Pep.\n\nContactjahorst@gmail.com, ram@compbio.org.\n\nSupplementary informationAvailable at Bioinformatics online.

bioinformatics

Rice protein models from the Nutritious Rice for the World Project

BackgroundMany rice protein sequences are very different from the sequence of proteins with known structures. Homology modeling is not possible for many rice proteins. However, it is possible to use computational intensive de novo techniques to obtain protein models when the protein sequence cannot be mapped to a protein of known structure. The Nutritious Rice for the World project generated 10 billion models encompassing more than 60,000 small proteins and protein domains for the rice strains Oryza sativa and Oryza japonica.\n\nFindingsOver a period of 1.5 years, the volunteers of World Community Grid supported by IBM generated 10 billion candidate structures, a task that would have taken a single CPU on the order of 10 millennia. For each protein sequence, 5 top structures were chosen using a novel clustering methodology developed for analyzing large datasets. These are provided along with the entire set of 10 billion conformers.\n\nConclusionsWe anticipate that the centroid models will be of use in visualizing and determining the role of rice proteins where the function is unknown. The entire set of conformers is unique in terms of size and that they were derived from sequences that lack detectable homologs. Large sets of de novo conformers are rare and we anticipate that this set will be useful for benchmarking and developing new protein structure prediction methodologies.

plant biology