bioRxiv ScienceSearch

Biology subjects

Verschaffelt, P.

Publications and source records attributed to Verschaffelt, P..

4 recordsLinked to original sources

Pout2Prot: an efficient tool to create protein (sub)groups from Percolator output files

In metaproteomics, the study of the collective proteome of microbial communities, the protein inference problem is more challenging than in single-species proteomics. Indeed, a peptide sequence can not only be present in multiple proteins or protein isoforms of the same species, but also in homologous proteins from closely related species. To assign the taxonomy and functions of the microbial species, specialized tools have been developed, such as Prophane. This tool, however, is not directly compatible with post-processing tools such as Percolator. In this manuscript we therefore present Pout2Prot, which takes Percolator Output (.pout) files from multiple experiments and creates protein group and protein subgroup output files (.tsv) that can be used directly with Prophane. We investigated different grouping strategies, and compared existing protein grouping tools to develop an advanced protein grouping algorithm that offers a variety of different approaches, allows grouping for multiple files, and uses a weighted spectral count for protein (sub)groups to reflect abundance. Pout2Prot is available as a web application at https://pout2prot.ugent.be and is installable via pip as a standalone command line tool and reusable software library. All code is open source under the Apache License 2.0 and is available at https://github.com/compomics/pout2prot.

bioinformatics

UMGAP: the Unipept MetaGenomics Analysis Pipeline

Shotgun metagenomics is now commonplace to gain insights into communities from diverse environments, but fast, memory-friendly, and accurate tools are needed for deep taxonomic analysis of the metagenome data. To meet this need we developed UMGAP, a highly versatile open source command line tool implemented in Rust for taxonomic profiling of shotgun metagenomes. It differs from state-of-the-art tools in its use of protein code regions identified in short reads for robust taxonomic identifications, a broad-spectrum index that can identify both archaea, bacteria, eukaryotes and viruses, a non-monolithic design, and support for interactive visualizations of complex biodiversities.

bioinformatics

Critical Assessment of Metaproteome Investigation (CAMPI): a Multi-Lab Comparison of Established Workflows

Metaproteomics has matured into a powerful tool to assess functional interactions in microbial communities. While many metaproteomic workflows are available, the impact of method choice on results remains unclear. Here, we carried out the first community-driven, multi-laboratory comparison in metaproteomics: the critical assessment of metaproteome investigation study (CAMPI). Based on well-established workflows, we evaluated the effect of sample preparation, mass spectrometry, and bioinformatic analysis using two samples: a simplified, laboratory-assembled human intestinal model and a human fecal sample. We observed that variability at the peptide level was predominantly due to sample processing workflows, with a smaller contribution of bioinformatic pipelines. These peptide-level differences largely disappeared at the protein group level. While differences were observed for predicted community composition, similar functional profiles were obtained across workflows. CAMPI demonstrates the robustness of present-day metaproteomics research, serves as a template for multi-laboratory studies in metaproteomics, and provides publicly available data sets for benchmarking future developments.

microbiology

MegaGO: a fast yet powerful approach to assess functional similarity across meta-omics data sets

The study of microbiomes has gained in importance over the past few years, and has led to the fields of metagenomics, metatranscriptomics and metaproteomics. While initially focused on the study of biodiversity within these communities the emphasis has increasingly shifted to the study of (changes in) the complete set of functions available in these communities. A key tool to study this functional complement of a microbiome is Gene Ontology (GO) term analysis. However, comparing large sets of GO terms is not an easy task due to the deeply branched nature of GO, which limits the utility of exact term matching. To solve this problem, we here present MegaGO, a user-friendly tool that relies on semantic similarity between GO terms to compute functional similarity between two data sets. MegaGO is highly performant: each set can contain thousands of GO terms, and results are calculated in a matter of seconds. MegaGO is available as a web application at https://megago.ugent.be and installable via pip as a standalone command line tool and reusable software library. All code is open source under the MIT license, and is available at https://github.com/MEGA-GO/.

bioinformatics