bioRxiv Science⌕ Search

Biology subjects

Malik, A. J.

Publications and source records attributed to Malik, A. J..

4 recordsLinked to original sources

On use of tertiary structure characters in hidden Markov models for protein fold prediction

While advances in protein structure prediction have opened up insights into arcane proteins, weak sequence homology makes functional characterisation challenging. To overcome this challenge, we use structure-based hidden Markov models of groupings in SCOP, CATH and ECOD to predict folds in proteins and thereby infer function. Conservation of structure and ability of hidden Markov models to detect remote signals make this a powerful resource for complete characterisation of arcane proteins.

bioinformatics↗

Tertiary-interaction characters enable fast, model-based structural phylogenetics beyond the twilight zone

Protein structure is more conserved than protein sequence, and therefore may be useful for phylogenetic inference beyond the "twilight zone" where sequence similarity is highly decayed. Until recently, structural phylogenetics was constrained by the lack of solved structures for most proteins, and the reliance on phylogenetic distance methods which made it difficult to treat inference and uncertainty statistically. AlphaFold has mostly overcome the first problem by making structural predictions readily available. We address the second problem by redeploying a structural alphabet recently developed for Foldseek, a highly-efficient deep homology search program. For each residue in a structure, Foldseek identifies a tertiary interaction closest-neighbor residue in the structure, and classifies it into one of twenty "3Di" states. We test the hypothesis that 3Dis can be used as standard phylogenetic characters using a dataset of 53 structures from the ferritin-like superfamily. We performed 60 IQtree Maximum Likelihood runs to compare structure-free, PDB, and AlphaFold analyses, and default versus custom model sets that include a 3DI-specific rate matrix. Analyses that combine amino acids, 3Di characters, partitioning, and custom models produce the closest match to the structural distances tree of Malik et al. (2020), avoiding the long-branch attraction errors of structure-free analyses. Analyses include standard ultrafast bootstrapping confidence measures, and take minutes instead of weeks to run on desktop computers. These results suggest that structural phylogenetics could soon be routine practice in protein phylogenetics, allowing the re-exploration of many fundamental phylogenetic problems.

evolutionary biology↗

On quantum computing and geometry optimization

Quantum computers have demonstrated advantage in tackling problems considered hard for classical computers and hold promise for tackling complex problems in molecular mechanics such as mapping the conformational landscapes of biomolecules. This work attempts to explore a few ways in which classical data, relating to the Cartesian space representation of biomolecules, can be encoded for interaction with empirical quantum circuits not demonstrating quantum advantage. Using the quantum circuit in a variational arrangement together with a classical optimizer, this work deals with the optimization of spatial geometries with potential application to molecular assemblies. Additionally this work uses quantum machine learning for protein side-chain rotamer classification and uses an empirical quantum circuit for random state generation for Monte Carlo simulation for side-chain conformation sampling. Altogether, this novel work suggests ways of bridging the gap between conventional problems in life sciences and how potential solutions can be obtained using quantum computers. It is hoped that this work will provide the necessary impetus for wide-scale adoption of quantum computing in life sciences.

bioinformatics↗

Structome: Exploring the structural neighbourhood of proteins

Protein structures carry signal of common ancestry and can therefore aid in reconstructing their evolutionary histories. To expedite the structure-informed inference process, a web server, Structome, has been developed, that allows users to rapidly identify protein structures similar to a query protein and to assemble datasets useful for structure-based phylogenetics. Structome was created by clustering[~] 94% of the structures in RCSB PDB using 90% sequence identity and representing each cluster by a centroid structure. Structure similarity between centroid proteins was calculated, and annotations from PDB, SCOP and CATH were integrated. To illustrate utility, an H3 histone was used as a query, and results show that the protein structures returned by Structome span both sequence and structural diversity of the histone fold. Additionally, the pre-computed nexus-formated distance matrix, provided by Structome, enables analysis of evolutionary relationships between proteins not identifiable using searches based on sequence similarity alone. Our results demonstrate that, beginning with a single structure, Structome can be used to rapidly generate a dataset of structural neighbours and allows deep evolutionary history of proteins to be studied. Structome is available at: https://structome.bii.a-star.edu.sg

bioinformatics↗