bioRxiv Science⌕ Search

Biology subjects

Tahmid, M. T.

Publications and source records attributed to Tahmid, M. T..

10 recordsLinked to original sources

GraFusionNet: Integrating Node, Edge, and Semantic Features for Enhanced Graph Representations

Understanding complex graph-structured data is a cornerstone of modern research in fields like cheminformatics and bioinformatics, where molecules and biological systems are naturally represented as graphs. However, traditional graph neural networks (GNNs) often fall short by focusing mainly on node features while overlooking the rich information encoded in edges. To bridge this gap, we present GraFusionNet, a framework designed to integrate node, edge, and molecular-level semantic features for enhanced graph classification. By employing a dual-graph autoencoder, GraFusionNet transforms edges into nodes via a line graph conversion, enabling it to capture intricate relationships within the graph structure. Additionally, the incorporation of Chem-BERT embeddings introduces semantic molecular insights, creating a comprehensive feature representation that combines structural and contextual information. Our experiments on benchmark datasets, such as Tox21 and HIV, highlight GraFusionNets superior performance in tasks like toxicity prediction, significantly surpassing traditional models. By providing a holistic approach to graph data analysis, GraFusion-Net sets a new standard in leveraging multi-dimensional features for complex predictive tasks. CCS CONCEPTSO_LIComputing methodologies [->] Neural networks. C_LI ACM Reference FormatMd Toki Tahmid, Tanjeem Azwad Zaman, and Mohammad Saifur Rahman. 2018. GraFusionNet: Integrating Node, Edge, and Semantic Features for Enhanced Graph Representations. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym XX). ACM, New York, NY, USA, 9 pages. https://doi.org/XXXXXXX.XXXXXXX

molecular biology↗

TomoPicker: Annotation-Efficient Particle Picking in cryo-electron Tomograms

MotivationLocalizing macromolecules in crowded cellular cryo-electron tomograms (cryo-ET) is crucial for determining their in situ structures. Traditional template matching-based approaches for this task suffer from template-specific biases and have low throughput. Given these problems, learning-based solutions are necessary. However, the paucity of annotated data for training poses substantial challenges for such learning-based methods. Moreover, preparing extensively annotated cellular cryo-ET tomograms for training macromolecule localization methods is extremely time-consuming and burdensome due to the large volume and low signal-to-noise ratio of the tomograms. ResultsIn this work, we developed TomoPicker, an annotation-efficient macromolecule localization method for cryo-ET tomograms. To achieve such annotation-efficiency, TomoPicker regards macromolecule localization as a voxel classification problem and solves it with two different positive-unlabeled learning approaches. We evaluated TomoPicker on two experimental cryo-electron tomography (cryo-ET) datasets of crowded eukaryotic cells and one experimental dataset of relatively less crowded prokaryotic cell. We observed that, with only 10 annotated macromolecule locations, TomoPicker with positive unlabeled learning achieved a performance comparable to that of state-of-the-art supervised methods trained with several hundred annotations. In other words, TomoPicker achieved plausible segmentation with up to 98% less data compared to supervised learning-based methods. Furthermore, it demonstrated substantial improvements over existing learning-based macromolecule localization methods under sparse annotation scenarios. CodeThe code to train and use TomoPicker is available on https://github.com/DuranRafid/TomoPicker.

bioinformatics↗

DeepRNA-Twist: Language Model guided RNA Torsion Angle Prediction with Attention-Inception Network

RNA torsion and pseudo-torsion angles are critical in determining the three-dimensional conformation of RNA molecules, which in turn governs their biological functions. However, current methods are limited by RNAs structural complexity and flexibility, as it can adopt multiple conformations, with experimental techniques being costly and computational approaches struggling to capture the intricate sequence dependencies needed for accurate predictions. To address these challenges, we introduce DeepRNA-Twist, a novel deep learning framework designed to predict RNA torsion and pseudo-torsion angles directly from sequence. DeepRNA-Twist utilizes RNA language model embeddings, which provides rich, context-aware feature representations of RNA sequences. Additionally, it introduces 2A3IDC module (Attention Augmented Inception Inside Inception with Dilated CNN), combining inception networks with dilated convolutions and multi-head attention mechanism. The dilated convolutions capture long-range dependencies in the sequence without requiring a large number of parameters, while the multi-head attention mechanism enhances the models ability to focus on both local and global structural features simultaneously. DeepRNA-Twist was rigorously evaluated on benchmark datasets, including RNA-Puzzles, CASP-RNA, and SPOT-RNA-1D, and demonstrated significant improvements over existing methods, achieving state-of-the-art accuracy. Source code is available at https://github.com/abrarrahmanabir/DeepRNA-Twist

bioinformatics↗

Advancing Noninvasive Mechanical Ventilation:Simulating Techniques for Improved Respiratory Care

Respiratory failure is a critical condition that often requires mechanical ventilation to support or restore normal breathing. While invasive mechanical ventilation (IMV) is commonly used for severe cases, noninvasive mechanical ventilation (NIMV) offers a less intrusive alternative that reduces complications and can be applied in moderate cases. The COVID-19 pandemic highlighted the global shortage of ventilators, particularly in low- and middle-income countries (LMICs), where limited access to life-saving equipment exacerbated the crisis. In response to these challenges, this paper presents a simplified, compartmental-based simulation model for NIMV. This model provides a practical and accessible tool for simulating respiratory system behavior under various ventilation modes, using the analogy between electrical circuits and lung physiology. By simulating key parameters such as airway resistance and lung compliance, the model allows clinicians and researchers to evaluate ventilator performance and optimize treatment strategies. Furthermore, the simulation offers a blueprint for developing cost-effective, easy-to-use NIMV systems that can be deployed in resource-constrained environments. Our contribution seeks to address the ventilator shortage by enabling better design and understanding of noninvasive ventilation, ultimately improving respiratory care for patients with moderate respiratory failure.

bioengineering↗

BioLLMNet: Enhancing RNA-Interaction Prediction with a Specialized Cross-LLM Transformation Network

Existing computational methods for the prediction of RNA related interactions often rely heavily on manually crafted features. Language model features for bio-sequences has gain significant popularity in proteomics and genomics. However, during interaction prediction, how language model features from different modalities should be combined to extract the most representative features is yet to be explored. We introduce BioLLMNet, a novel framework that introduces an effective combination approach for multi-modal bio-sequences. BioLLMNet provides a way to transform feature space of different molecules language model features and uses learnable gating mechanism to effectively fuse features. Rigorous evaluations show that BioLLMNet achieves state-of-the-art performance in RNA-protein, RNA-small molecule, and RNA-RNA interactions, outperforming existing methods in RNA-associated interaction prediction.

bioinformatics↗

EmbedSimScore: Advancing Protein Similarity Analysis with Structural and Contextual Embeddings

Accurately computing protein similarity is challenging due to the intricate interplay between local substructures and the global structure within protein molecules. Traditional metrics like TM-score often focus on aligning the global structures of the proteins in a rather geometry-based algorithmic way, potentially overlooking critical local-global relations and contextual comparisons. We introduce Embed-SimScore, a novel self-supervised method that generates structural and contextual embeddings by jointly considering both local substructures and global proteins structures. Utilizing contrastive language-structure pre-training (CLSP) and structural contrastive learning, EmbedSimScore captures comprehensive features across different scales of protein structure. These embeddings provide a more precise and holistic means of computing protein similarities, resulting in the identification of intrinsic relations among proteins that traditional approaches overlook.

bioinformatics↗

LOCAS: Multi-label mRNA Localization with Supervised Contrastive Learning

Traditional methods for mRNA subcellular localization often fail to account for multiple compartmentalization. Recent multi-label models have improved performance, but still face challenges in capturing complex localization patterns. We introduce LOCAS (Localization with Supervised Contrastive Learning), which integrates an RNA language model to generate initial embeddings, employs supervised contrastive learning (SCL) to identify distinct RNA clusters, and uses a multi-label classification head (ML-Decoder) with cross-attention for accurate predictions. Through extensive ablation studies and multi-label overlapping threshold tuning, LOCAS achieves state-of-the-art performance across all metrics, providing a robust solution for RNA localization tasks.

bioinformatics↗

RNA-DCGen: Dual Constrained RNA Sequence Generation with LLM-Attack

Designing RNA sequences with specific properties is critical for developing personalized medications and therapeutics. While recent diffusion and flow-matching-based generative models have made strides in conditional sequence design, they face two key limitations: specialization for fixed constraint types, such as tertiary structures, and lack of flexibility in imposing additional conditions beyond the primary property of interest. To address these challenges, we introduce RNA-DCGen, a generalized framework for RNA sequence generation that is adaptable to any structural or functional properties through straightforward finetuning with an RNA language model (RNA-LM). Additionally, RNA-DCGen can enforce conditions on the generated sequences by fixing specific conserved regions. On RNA generation conditioned on RNA distance maps, RNA-DCGen generates sequences with an average R2 score of 0.625 compared to random sequences that score only 0.118 over 250 generations as judged by a separate more capable RNA-LM. When conditioned on RNA secondary structures, RNA-DCGen achieves an average F1 score of 0.4 against a random baseline of 0.006.

bioinformatics↗

wQFM-TREE: highly accurate and scalable quartet-based species tree inference from gene trees

Summary methods are becoming increasingly popular for species tree estimation from multi-locus data in the presence of gene tree discordance. ASTRAL, a leading method in this class, solves the Maximum Quartet Support Species Tree problem within a constrained solution space constructed from the input gene trees. In contrast, alternative heuristics such as wQFM and wQMC operate by taking a set of weighted quartets as input and employ a divide-and-conquer strategy to construct the species tree. Recent studies showed wQFM to be more accurate than ASTRAL and wQMC, though its scalability is hindered by the computational demands of explicitly generating and weighting {Theta}(n4) quartets. Here, we introduce wQFM-TREE, a novel summary method that enhances wQFM by circumventing the need for explicit quartet generation and weighting, thereby enabling its application to large datasets. Unlike wQFM, wQFM-TREE can also handle polytomies. Extensive simulations under diverse and challenging model conditions, with hundreds or thousands of taxa and genes, consistently demonstrate that wQFM-TREE matches or improves upon the accuracy of ASTRAL. Specifically, wQFM-TREE outperformed ASTRAL in 25 of 27 model conditions analyzed in this study involving 200-1000 taxa, with statistically significant differences in 20 of these conditions. Moreover, we applied wQFM-TREE to re-analyze the green plant dataset from the One Thousand Plant Transcriptomes Initiative. Its remarkable accuracy and scalability position wQFM-TREE as a highly competitive alternative to leading methods in the field. Additionally, the algorithmic and combinatorial innovations introduced in this study will benefit various quartet-based computations, advancing the state-of-the-art in phylogenetic estimations.

evolutionary biology↗

BiRNA-BERT Allows Efficient RNA Language Modeling with Adaptive Tokenization

Recent advancements in Transformer-based models have spurred interest in their use for biological sequence analysis. However, adapting models like BERT is challenging due to sequence length, often requiring truncation for proteomics and genomics tasks. Additionally, advanced tokenization and relative positional encoding techniques for long contexts in NLP are often not directly transferable to DNA/RNA sequences, which require nucleotide or character-level encodings for tasks such as 3D torsion angle prediction. To tackle these challenges, we propose an adaptive dual tokenization scheme for bioinformatics that utilizes both nucleotide-level (NUC) and efficient BPE tokenizations. Building on the dual tokenization, we introduce BiRNA-BERT, a 117M parameter Transformer encoder pretrained with our proposed tokenization on 28 billion nucleotides across 36 million coding and non-coding RNA sequences. The learned representation by BiRNA-BERT generalizes across a range of applications and achieves state-of-the-art results in long-sequence downstream tasks and achieves a performance comparable to 6x larger models in short-sequence tasks with 27xless pre-training compute. BiRNA-BERT can dynamically adjust its tokenization strategy based on sequence lengths, utilizing NUC for shorter sequences and switching to BPE for longer ones, thereby offering, for the first time, the capability to efficiently handle arbitrarily long DNA/RNA sequences. 1

bioinformatics↗