bioRxiv Science⌕ Search

Biology subjects

Joshi, J. P.

Publications and source records attributed to Joshi, J. P..

2 recordsLinked to original sources

Large-scale classification of metagenomic samples: a comparative analysis of classical machine learning techniques vs a novel brain-inspired hyperdimensional computing approach

Classical machine learning techniques have revolutionized bioinformatics, enabling researchers to extract knowledge from complex biological data. However, these techniques often struggle with high-dimensional data, where the increasing number of features leads to decreased performance, also affecting models accuracy. To address this problem, we explore hyperdimensional computing (HDC), an emerging brain-inspired computational paradigm that leverages high-dimensional vectors and simple arithmetic operations to represent and manipulate complex patterns, as an alternative approach in the context of supervised machine learning. In this work, we present a comprehensive comparative analysis of HDC against established machine learning techniques across a range of classification tasks. As a representative use case, we focus on classifying heterogeneous metagenomic samples based on their quantitative microbial profiles, using publicly available microbiome datasets. Our results demonstrate that HDC achieves comparable, and in some cases, superior classification accuracy to classical methods. Furthermore, our findings highlight the potential of HDC for improved computational efficiency, particularly when dealing with large-scale datasets, suggesting the HDC-based classifier as a promising tool for bioinformatics research, particularly in areas characterized by high-dimensional data. We also offer a Galaxy powered toolset to analyze your own datasets and generate reproducible workflows and adopt these methods in your own research with ease. Our investigation into the application of a HDC-based supervised machine learning technique for classifying microbial profiles in metagenomic samples yielded promising results, demonstrating the potential of this novel computational paradigm to complement and, in some cases, surpass the performances of well established machine learning techniques. ImportanceThe growing complexity and dimensionality of biological data require more efficient and scalable machine learning approaches. HDC offers a novel alternative to conventional methods, showing resilience to high-dimensionality while maintaining competitive accuracy. This study demonstrates the effectiveness of HDC in classifying metagenomic samples based on their microbial composition. Our results suggest that HDC not only matches, but sometimes exceeds the performance of well-established methods. We make this approach accessible to the broader bioinformatics community with an open-source tool fully integrated into the Galaxy platform, facilitating its adoption and reproducibility, with the aim of integrating HDC into mainstream biological data analysis pipelines, especially for complex, high-dimensional tasks in microbiome research.

microbiology↗

Integrated forward and reverse degradomics uncovers the proteolytic landscape of aortic aneurysms and the roles of MMP9 and mast cell chymase

BackgroundDysregulated proteolysis is implicated in thoracic (TAA) and abdominal aortic aneurysm (AAA) pathogenesis, but the proteolytic landscapes (degradomes) of aneurysmal and normal aorta, and contributions of individual proteases remain undefined. Here, a proteome-wide approach was used to uncover TAA and AAA degradomes, compare them quantitatively and define the specific role in aortic remodeling of two proteases consistently identified in the aneurysms, mast cell chymase (CMA1) and matrix metalloprotease 9 (MMP9). MethodsThe mass spectrometry-based N-terminomics strategy Terminal Amine Isotopic Labeling of Substrates (TAILS) was applied to Marfan syndrome TAAs (n=5), AAAs (n=16) and corresponding non-diseased aorta (TAs, n=4, and AAs, n=8) as a forward degradomics application, i.e., to define substrate and protease degradomes, and 8-plex iTRAQ-TAILS was used for quantitative comparison. Cleavage sites of CMA1 and MMP9 were sought by reverse degradomics, i.e., digestion of aortic proteins with these proteases, followed by 6-plex iTRAQ-TAILS. CMA1 and MMP9 proteolysis of biglycan was investigated using Amino-Terminal Oriented Mass spectrometry of Substrates (ATOMS). ResultsWe experimentally annotated 16,923 proteolytically derived peptides (substrate degradome) and 90 proteases (protease degradome) in the aorta. Quantitative substrate degradome comparisons identified specific differentially modulated pathways and networks in TAA and AAA. Reverse degradomics elucidated > 300 CMA1 and MMP9 substrate cleavage sites, of which, many, including orthogonally validated biglycan cleavage, occurred in the disease degradomes. ConclusionsUnbiased, proteome-wide forward degradomics of the aortic wall from TAA, AAA and non-diseased tissue generated the first systems biology view of vascular wall breakdown and public resource for the hitherto occult proteolytic landscape, demonstrating widespread extracellular matrix remodeling. The findings provide insights on aortic aneurysm pathways and potential disease biomarkers. Mapping of specific contributions of CMA1 and MMP9 on the aortic forward substrate degradome using reverse degradomics provides a strategy for defining the activities of all proteases involved in aortic disease.

biochemistry↗