bioRxiv Science⌕ Search

Biology subjects

Dalkiran, A.

Publications and source records attributed to Dalkiran, A..

4 recordsLinked to original sources

On the state of protein function prediction: a report on the fourth CAFA challenge

BackgroundThe Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). ResultsCAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

bioinformatics↗

Transfer Learning Enables Drug-Target Interaction Prediction in Data-Scarce One-Carbon Metabolism

Predicting drug-target interactions (DTIs) with deep learning offers opportunities to accelerate drug discovery, yet performance is constrained by the scarcity of target-specific training data. This is a particular challenge for mitochondrial one-carbon (1C) pathway enzymes, which are attractive therapeutic targets but remain pharmacologically understudied. Mitochondrial 1C metabolism supplies glycine, reducing equivalents, and 1C units critical for nucleotide synthesis, and has emerged as a key pathway in cancer and fibrosis. SHMT2 and MTHFD2, two key 1C enzymes, support collagen production in fibroblasts, blocking either prevents TGF-{beta}-induced glycine and collagen accumulation. Here, we developed transfer learning-based deep learning models to predict interactions between approved drugs and SHMT2 or MTHFD2 despite minimal target-specific training data, pre-training on large datasets from related enzymes before fine-tuning to these targets. Virtual screening of the DrugBank library identified six candidates, three of which, Carbimazole, Crizotinib, and GSK2018682 reduced TGF-{beta}-induced collagen production and glycine accumulation in human lung fibroblasts, demonstrating transfer learning as a strategy for repurposable drug identification in data-scarce metabolic targets.

bioinformatics↗

Molecular Contrastive Learning with Graph Attention Network (MoCL-GAT) for Enhanced Molecular Representation

Learning the representation of molecules is crucial for drug discovery but is often hindered by the scarcity of labeled experimental data, which limits the performance of supervised machine learning models. While self-supervised learning (SSL) offers a solution by leveraging vast unlabeled chemical databases, many existing methods focus on learning from either local structural information or global molecular properties, but not both simultaneously. We introduce MoCL-GAT, a novel contrastive and transfer learning-based SSL framework that addresses this gap by simultaneously learning from two complementary objectives. It combines a local contrastive task on molecular subgraphs to capture fine-grained chemical environments with a global predictive task to learn holistic molecular descriptors. This dual-objective approach, powered by a Graph Attention Network, is designed to create more robust, versatile, and transferable molecular representations. Pre-trained on 1.9 million compounds, MoCL-GAT was fine-tuned on diverse benchmarks. It achieved state-of-the-art performance on molecular property prediction tasks, with an AUROC of 0.928 on BBBP and 0.749 on SIDER, and top-ranking RMSEs of 0.570 for ESOL and 1.818 for FreeSolv. Critically, fine-tuned models consistently and significantly outperformed models trained from scratch, confirming the value of pre-training. These results validate that MoCL-GATs dual-objective approach learns highly effective and transferable representations, enabling more accurate and data-efficient predictions for key cheminformatics challenges.

bioinformatics↗

Discovery of Electron Hole-hopping Redox Mutations in Myoglobin by Deep Mutational Learning

In addition to storing molecular oxygen, myoglobin catalyzes peroxidase-like reactions involving high valency iron(IV)-oxo species that support one-electron oxidations on a range of substrates at an open active site. In select metalloenzymes, long-range electron transfer can be mediated by hole-hopping pathways composed of aromatic residues that act as relay stations for oxidative equivalents. However, it remains unclear how sequence variations could introduce or alter such catalytic mechanisms in myoglobin. Here we used enzyme proximity sequencing (EP-Seq) to measure the peroxidase activity levels of >6,000 human myoglobin variants. The resulting fitness landscape reveals how aromatic substitutions, in particular surface-exposed tryptophans, can enhance peroxidase activity. Using protein language models in tandem with feedforward neural networks, we trained an accurate fitness predictor on the experimental dataset, and applied it to screen >4M double mutant variants. The predictions suggested a beneficial role for electron-hole-hopping mutations in improving peroxidase activity. We experimentally tested 20 high scoring variants in a yeast display assay, all of which outperformed wild type myoglobin. Three selected variants were also tested in soluble format and similarly showed improved performance. A focused combinatorial library yielded a top double tryptophan variant (Q92W/F107W) with 4.9-fold higher catalytic efficiency than wild type. These results show that deep mutational learning can identify myoglobin variants with enhanced peroxidase activity that are consistent with the involvement of hole-hopping pathways, with broad implications for biocatalyst and redox enzyme design.

bioengineering↗