bioRxiv Science⌕ Search

Biology subjects

Arefian, M.

Publications and source records attributed to Arefian, M..

3 recordsLinked to original sources

Spectronaut-nf: A Nextflow Pipeline for Parallel Processing of DIA Data with Spectronaut

SummaryContemporary proteomics methods can now generate large-scale DIA datasets of thousands of files that demand substantial computational resources for efficient analysis. Spectronaut is a widely used platform for DIA data processing; however, large-scale searches are often constrained by computational performance and long execution times when run on single workstations. Here, we present Spectronaut-nf, a Nextflow-based pipeline that enables scalable and parallelized execution of Spectronaut analyses across high-performance computing (HPC) environments. The workflow divides directDIA analysis into modular stages, including spectral library generation, DIA searching, and merging results, allowing efficient distribution of tasks across multiple compute nodes. Benchmarking using 72 diaPASEF raw files using typical hardware demonstrated that Spectronaut-nf completed searches in 23.77 hours, compared with 39.09 hours on a Windows workstation and 67.04 hours on a single-node Linux HPC setup. Stress testing with 1,037 diaPASEF raw files further demonstrated the scalability and robustness of the workflow for large proteomics datasets. Across platforms, protein and peptide identifications remained consistent, with only minimal variability attributable to platform-specific differences. Overall, Spectronaut-nf provides a flexible, scalable, and efficient framework for high-throughput DIA proteomics analysis in HPC environments. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=129 SRC="FIGDIR/small/741433v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@4a7107org.highwire.dtl.DTLVardef@14287d8org.highwire.dtl.DTLVardef@e484e9org.highwire.dtl.DTLVardef@d21c07_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Proteogenomic discovery of novel small proteins in clinical Mycobacterium tuberculosis strains

Even though our meta-analysis ranks Mycobacterium tuberculosis genomes among the bacterial pathogens that are most straightforward to assemble, most available assemblies relied on short-read sequencing and contain genomic blind spots that miss functionally important genes. Complete genomes are essential for functional genomics, particularly for identifying small ORF-encoded proteins (SEPs; [≤]100 amino acids), which can play critical biological roles yet are frequently missed by standard annotations. Here, we generated complete long-read assemblies for six clinical reference strains representing lineage 1 and the more pathogenic lineage 2, followed by comparative genomic and proteogenomic analyses. We additionally provide software to predict comprehensive sets of mycobacteria-specific proline-glutamic acid (PE) and PPE family genes, including lineage-specific variants. Using parallel accumulation-serial fragmentation mass spectrometry, we detected approximately two-thirds of each strains annotated proteome from unfractionated cell extracts. Extending our proteogenomic framework across related strains, and adding rigorous control of proteogenomic discovery rates using entrapment strategies, we revealed 12-24 previously unannotated proteins per strain, predominantly SEPs, 56-60 alternative translation start sites, and 9-17 expressed pseudogenes. Newly identified proteins included conserved and lineage-specific SEPs, an antitoxin, candidate antimicrobial peptides and novel proteins under purifying selection. Overall, applying this improved proteogenomics method to phylogenomically selected clinical reference strains provides a valuable approach for discovering candidate diagnostics or therapeutics, as illustrated here for a WHO-listed critical bacterial pathogen.

microbiology↗

Induced pluripotent stem cell-derived macrophages as a model for human inflammasome signaling

Macrophage models are a mainstay of inflammasome research, however current human in vitro macrophage models have significant limitations. Here we generate induced pluripotent stem cell (iPSC)-derived macrophages (iMacs) to study inflammasome signaling and benchmark them with human monocyte-derived macrophages (HMDMs). We confirm that iMacs express high levels of macrophage markers and are highly phagocytic. Whole cell proteomics analysis shows that iMacs express many inflammasome sensors and related proteins, and in functional assays iMacs respond to multiple inflammasome stimuli. The NLRP3 inflammasome is strongly activated in iMacs and we find that nigericin alone activates NLRP3. The non-canonical inflammasome does not require a priming step in iMacs as caspase-4 is constitutively expressed. High levels of NAIP/NLRC4 inflammasome activation are also observed in response to needle toxin. Finally, unlike HMDMs, iMacs activate NLRP1. Therefore, we demonstrate that iMacs are a physiologically relevant and highly tractable model to study human inflammasome signaling and regulation. MotivationiPSC-derived macrophages (iMacs) are functionally, transcriptionally, and phenotypically similar to primary human macrophages. iMacs therefore offer new opportunities to study inflammasome activity in a human macrophage model, but to date they have not been widely used. In this study, we describe a protocol to differentiate and characterize iMacs. We then describe how to activate a range of different inflammasomes within these cells and assess the inflammasome response by measuring pyroptosis, cytokine release, ASC speck formation, and processing of inflammasome-related proteins. We also benchmark iMac responses with the current gold standard primary human monocyte derived macrophage model.

immunology↗