bioRxiv Science⌕ Search

Biology subjects

Dokholyan, N.

Publications and source records attributed to Dokholyan, N..

3 recordsLinked to original sources

Application of Quantum Tensor Networks for Protein Classification

Computational methods in drug discovery significantly reduce both time and experimental costs. Nonetheless, certain computational tasks in drug discovery can be daunting with classical computing techniques which can be potentially overcome using quantum computing. A crucial task within this domain involves the functional classification of proteins. However, a challenge lies in adequately representing lengthy protein sequences given the limited number of qubits available in existing noisy quantum computers. We show that protein sequences can be thought of as sentences in natural language processing and can be parsed using the existing Quantum Natural Language framework into parameterized quantum circuits of reasonable qubits, which can be trained to solve various proteinrelated machine-learning problems. We classify proteins based on their sub-cellular locations--a pivotal task in bioinformatics that is key to understanding biological processes and disease mechanisms. Leveraging the quantum-enhanced processing capabilities, we demonstrate that Quantum Tensor Networks (QTN) can effectively handle the complexity and diversity of protein sequences. We present a detailed methodology that adapts QTN architectures to the nuanced requirements of protein data, supported by comprehensive experimental results. We demonstrate two distinct QTNs, inspired by classical recurrent neural networks (RNN) and convolutional neural networks (CNN), to solve the binary classification task mentioned above. Our top-performing quantum model has achieved a 94% accuracy rate, which is comparable to the performance of a classical model that uses the ESM2 protein language model embeddings. Its noteworthy that the ESM2 model is extremely large, containing 8 million parameters in its smallest configuration, whereas our best quantum model requires only around 800 parameters. We demonstrate that these hybrid models exhibit promising performance, showcasing their potential to compete with classical models of similar complexity.

bioinformatics↗

Identification of post-transcriptional modifications in nucleic acid sequences using propose-designed molecular beacons

Post-transcriptional RNA modifications (PTxMs) present in small RNA species, specifically circulating extracellular RNAs, were recently identified as clinically relevant readouts, often more indicative of disease severity than the classical "up and down" changes in their copy number alone. While identification of PTxMs requires multiple and complex sample preparation steps, microgram-range amounts of RNA, followed by expensive and protracted bioinformatics analyses, the clinically relevant information is usually a yes/no for a particular genetic variant(s), and an up/down answer for relevant biomarkers. We have previously shown that molecular beacons (MBs) can identify specific nucleic acid sequences with picomolar sensitivity and single nucleotide specificity by exploiting the target-dependent change in their electrophoretic mobility profile. We now present a method for direct identification of miRNAs and isomiRs in cells and extracellular vesicles using gel electrophoresis, without the need for RNA isolation and purification. The detection is based on discreet changes in the hydrodynamic surface profile, the overall size, charge and charge distribution of the MB-target hybrid. Furthermore, using an RNA tertiary structure prediction algorithm (iFoldRNA) and a custom molecular dynamics simulation (DMD), we designed modified MBs specific for m6A-modified nucleotides in target RNA sequences. The sample preparation method coupled to the software package affords the design of specific MBs and sensitive, multiplex-type detection of targets in a wide variety of biofluids and cells, in a simple mix and read approach.

molecular biology↗

Thermodynamic impacts of combinatorial mutagenesis on protein conformational stability: precise, high-throughput measurement by Thermofluor

The D1 switch is a packing motif, broadly distributed in the proteome, that couples tryptophanyl-tRNA synthetase (TrpRS) domain movement to catalysis and specificity, thereby creating an escapement mechanism essential to free-energy transduction. The escapement mechanism arose from analysis of an extensive set of combinatorial mutations to this motif, which allowed us to relate mutant-induced changes quantitatively to both kinetic and computational parameters during catalysis. To further characterize the origins of this escapement mechanism in differential TrpRS conformational stabilities, we use high-throughput Thermofluor measurements for the 16 variants to extend analysis of the mutated residues to their impact on unliganded TrpRS stability. Aggregation of denatured proteins complicates thermodynamic interpretations of denaturation experiments. The free energy landscape of a liganded TrpRS complex, carried out for different purposes, closely matches the volume, helix content, and transition temperatures of Thermoflour and CD melting profiles. Regression analysis using the combinatorial design matrix accounts for >90% of the variance in Tms of both Thermofluor and CD melting profiles. We argue that the agreement of experimental melting temperatures with both computational free energy landscape and with Regression modeling means that experimental melting profiles can be used to analyze the thermodynamic impact of combinatorial mutations. Tertiary packing and aromatic stacking of Phenylalanine 37 exerts a dominant stabilizing effect on both native and molten globular states. The TrpRS Urzyme structure remains essentially intact at the highest temperatures explored by the simulations.

biophysics↗