bioRxiv Science⌕ Search

Biology subjects

Balasubramaniyan, B.

Publications and source records attributed to Balasubramaniyan, B..

2 recordsLinked to original sources

Large-scale analysis of ligand binding mode similarities in the PDB using interaction fingerprints

Three-dimensional structures of protein-ligand complexes are essential for insights into the molecular principles that govern ligand recognition and binding. With more than 180,000 ligand-bound entries in the Protein Data Bank (PDB), representing over two million individual complexes, the volume of available structural data offers unprecedented opportunities for large-scale analysis of interaction patterns. Analysis of interaction patterns across the PDB archive can help discover similarities and differences in the binding modes of ligands, assisting in drug discovery. However, large-scale analysis of up-to-date information remains a significant challenge due to the rapid growth of data. Here, we introduce the Extended Connectivity Interaction Fingerprint (ECIFP), an interaction-based fingerprint that simplifies 3D protein-ligand contact information into a fingerprint, while retaining key molecular and chemical features of the interacting fragments. The simpler fingerprint representation of the interaction data makes comparison of millions of protein-ligand complexes tractable. Benchmarking shows that ECIFP outperforms ligand-only Extended Connectivity Fingerprints in identifying similar binding sites across identical protein sequences occupied by chemically diverse ligands. Our analysis showed that similarities calculated using ECIFP can be used to compare macromolecular complexes with similar or different ligands. In this study, we demonstrate two large-scale applications of ECIFP: (1) identification of distinct binding modes for over 9,000 ligands across the entire PDB, and (2) detection of binding-mode similarities among structurally diverse ligands within the same binding site across 48,870 binding sites from over 21,000 proteins.

bioinformatics↗

Linking protein residues in literature and structure

Protein structures are crucial in understanding function, mechanism and disease-causing variants of proteins within any living cell. A number of experimental techniques are employed by researchers to determine said structure. Through structure inspection in molecular viewers combined with supporting biochemical and biophysical experiments, scientists are able to identify a proteins function, reaction mechanism and effects caused by sequence variation. These detailed findings supported by experimental results are documented and described in detail in scientific literature and by open sourcing the accompanying data. By writing a detailed report about the findings and providing evidence in additional files and complementary data formats it has become increasingly difficult for a reader, in particular a non-expert, to access the correct additional information and assess the validity of the drawn conclusion based on experimental results. It often requires a reader to resort to a number of different software packages to access the different data types. Here, we present a first-of-its-kind implementation of an artificial intelligence and text mining supported software tool that allows linking of text mentions of specific protein residues to their corresponding counterpart in the respective protein structure. An identified residue is highlighted in the publication text and upon interaction with the annotation, a molecule viewer displays the associated protein structures in the publication which contain said residue. The viewer is complemented by a display table that contains protein structure quality metrics for each occurrence of a residue. As such a reader can now explore a residue of interest they are currently assessing in a publication within its respective protein structure supported by its experimental evidence in a single view and application.

bioinformatics↗