bioRxiv Science⌕ Search

Biology subjects

Momenzadeh, A.

Publications and source records attributed to Momenzadeh, A..

3 recordsLinked to original sources

Machine Learning Identifies Plasma Proteomic Signatures of Descending Thoracic Aortic Disease

BackgroundDescending thoracic aortic aneurysms and dissections can go undetected until severe and catastrophic, and few clinical indices exist to screen for aneurysms or predict risk of dissection. MethodsThis study generated a plasma proteomic dataset from 75 patients with descending type B dissection (Type B) and 62 patients with descending thoracic aortic aneurysm (DTAA). Standard statistical approaches were compared to supervised machine learning (ML) algorithms to distinguish Type B from DTAA cases. Quantitatively similar proteins were clustered based on linkage distance from hierarchical clustering and ML models were trained with uncorrelated protein lists across various linkage distances with hyperparameter optimization using 5-fold cross validation. Permutation importance (PI) was used for ranking the most important predictor proteins of ML classification between disease states and the proteins among the top 10 PI protein groups were submitted for pathway analysis. ResultsOf the 1,549 peptides and 198 proteins used in this study, no peptides and only one protein, hemopexin (HPX), were significantly different at an adjusted p-value <0.01 between Type B and DTAA cases. The highest performing model on the training set (Support Vector Classifier) and its corresponding linkage distance (0.5) were used for evaluation of the test set, yielding a precision-recall area under the curve of 0.7 to classify between Type B from DTAA cases. The five proteins with the highest PI scores were immunoglobulin heavy variable 6-1 (IGHV6-1), lecithin-cholesterol acyltransferase (LCAT), coagulation factor 12 (F12), HPX, and immunoglobulin heavy variable 4-4 (IGHV4-4). All proteins from the top 10 most important correlated groups generated the following significantly enriched pathways in the plasma of Type B versus DTAA patients: complement activation, humoral immune response, and blood coagulation. ConclusionsWe conclude that ML may be useful in differentiating the plasma proteome of highly similar disease states that would otherwise not be distinguishable using statistics, and, in such cases, ML may enable prioritizing important proteins for model prediction.

systems biology↗

Complete Workflow for High Throughput Human Single Skeletal Muscle Fiber Proteomics

Skeletal muscle is a major regulatory tissue of whole-body metabolism and is composed of a diverse mixture of cell (fiber) types. Aging and several diseases differentially affect the various fiber types, and therefore, investigating the changes in the proteome in a fiber-type specific manner is essential. Recent breakthroughs in isolated single muscle fiber proteomics have started to reveal heterogeneity among fibers. However, existing procedures are slow and laborious requiring two hours of mass spectrometry time per single muscle fiber; 50 fibers would take approximately four days to analyze. Thus, to capture the high variability in fibers both within and between individuals requires advancements in high throughput single muscle fiber proteomics. Here we use a single cell proteomics method to enable quantification of single muscle fiber proteomes in 15 minutes total instrument time. As proof of concept, we present data from 53 isolated skeletal muscle fibers obtained from two healthy individuals analyzed in 13.25 hours. Adapting single cell data analysis techniques to integrate the data, we can reliably separate type 1 and 2A fibers. Sixty-five proteins were statistically different between clusters indicating alteration of proteins involved in fatty acid oxidation, muscle structure and regulation. Our results indicate that this method is significantly faster than prior single fiber methods in both data collection and sample preparation while maintaining sufficient proteome depth. We anticipate this assay will enable future studies of single muscle fibers across hundreds of individuals, which has not been possible previously due to limitations in throughput.

biochemistry↗

Parallelization with Dual-Trap Single-Column Configuration Maximizes Throughput of Proteomic Analysis

Proteomic analysis on the scale that captures population and biological heterogeneity over hundreds to thousands of samples requires rapid mass spectrometry methods which maximize instrument utilization (IU) and proteome coverage while maintaining precise and reproducible quantification. To achieve this, a short liquid chromatography gradient paired to rapid mass spectrometry data acquisition can be used to reproducibly profile a moderate set of analytes. High throughput profiling at a limited depth is becoming an increasingly utilized strategy for tackling large sample sets but the time spent on loading the sample, flushing the column(s), and re-equilibrating the system reduces the ratio of meaningful data acquired to total operation time and IU. The dual-trap single-column configuration presented here maximizes IU in rapid analysis (15 min per sample) of blood and cell lysates by parallelizing trap column cleaning and sample loading and desalting with analysis of the previous sample. We achieved 90% IU in low micro-flow (9.5 {micro}L/min) analysis of blood while reproducibly quantifying 300-400 proteins and over 6,000 precursor ions. The same IU was achieved for cell lysates, in which over 4,000 proteins (3,000 at CV below 20%) and 40,000 precursor ions were quantified at a rate of 15 minutes/sample. Thus, deployment of this dual-trap single column configuration enables high throughput epidemiological blood-based biomarker cohort studies and cell-based perturbation screening.

biochemistry↗