bioRxiv Science⌕ Search

Biology subjects

Fertin, G.

Publications and source records attributed to Fertin, G..

3 recordsLinked to original sources

Fast alignment of mass spectra in large proteomics datasets, capturing dissimilarities arising from multiple complex modifications of peptides

BackgroundIn proteomics, the interpretation of mass spectra representing peptides carrying multiple complex modifications is still challenging, currently limited by the number of potential modifications considered in a single analysis and the need to know them in advance. Further developments must be done in the field to help the scientific community to discover new post-translational modifications that play an essential role in disease and to understand how chemical modifications carried by food proteins could impact our health. ResultsTo make progress on this issue, we implemented SpecGlobX (SpecGlob eXTended to eXperimental spectra), a standalone Java application that quickly determines the best spectral alignments of a (possibly very large) list of Peptide-to-Spectrum Matches (PSMs) provided by any open modification search method, or generated by the user. As input, SpecGlobX reads a file containing spectra in MGF or mzML format and a semicolon-delimited spreadsheet describing the PSMs. As output, SpecGlobX returns the best alignment for each PSM, splitting the mass difference between the spectrum and the peptide into one or more shifts while considering the possibility of non-aligned masses (a phenomenon resulting from many situations including neutral losses). SpecGlobX is fast, able to align one million PSMs in about 1.5 minutes on a standard desktop. Firstly, we remind the foundations of the algorithm and detail how we adapted SpecGlob (the method we previously developed following the same aim, but limited to the interpretation of perfect simulated spectra) to the interpretation of imperfect experimental spectra. Then, we highlight the interest of SpecGlobX as a complementary tool downstream to three open modification search methods on a large simulated spectra dataset. Finally, we show on a smaller dataset that SpecGlobX performs equally well on experimental and simulated spectra. ConclusionsSpecGlobX is helpful as a decision support tool, providing keys to interpret peptides carrying complex modifications still poorly considered by current open modification search software. Better alignment of PSMs enhances confidence in the identification of spectra provided by open modification search methods and should improve the interpretation rate of spectra.

bioinformatics↗

SpecGlob: rapid and accurate alignment of mass spectra differing from their peptide models by several unknown modifications

BackgroundIn proteomics, mass spectra representing peptides carrying multiple unknown modifications are particularly difficult to interpret, which results in a large number of unidentified spectra. MethodsWe developed SpecGlob, a dynamic programming algorithm that aligns pairs of spectra, each such pair being a Peptide-Spectrum Match (PSM) provided by any Open Modification Search (OMS) method. For each PSM, SpecGlob computes the best alignment according to a given score system, while interpreting the mass delta within the PSM as one or several unspecified modification(s). All alignments are provided in a file, written in a specific syntax. ResultsUsing several sets of simulated spectra generated from the human proteome, we demonstrate that running SpecGlob as a post-analysis of an OMS method can significantly increase the number of correctly interpreted spectra, as SpecGlob is able to correctly and rapidly align spectra that differ by one or more modification(s) without any a priori. ConclusionSince SpecGlob explores all possible alignments that may explain the mass delta within a PSM, it reduces interpretation errors generated by incorrect assumptions about the modifications present in the sample or the number and the specificities of modifications carried by peptides. Our results demonstrate that SpecGlob should be relevant to align experimental spectra, although this consists in a more challenging task.

bioinformatics↗

MAGNETO: an automated workflow for genome-resolved metagenomics

Metagenome-Assembled Genomes (MAGs) represent individual genomes recovered from metagenomic data. MAGs are extremely useful to analyse uncultured microbial genomic diversity, as well as to characterize associated functional and metabolic potential in natural environments. Recent computational developments have considerably improved MAGs reconstruction but also emphasized several limitations, such as the non-binning of sequence regions with repetitions or distinct nucleotidic composition. Different assembly and binning strategies are often used, however, it still remains unclear which assembly strategy in combination with which binning approach, offers the best performance for MAGs recovery. Several workflows have been proposed in order to reconstruct MAGs, but users are usually limited to single-metagenome assembly or need to manually define sets of metagenomes to co-assemble prior to genome binning. Here, we present MAGNETO, an automated workflow dedicated to MAGs reconstruction, which includes a fully-automated co-assembly step informed by optimal clustering of metagenomic distances, and implements complementary genome binning strategies, for improving MAGs recovery. MAGNETO is implemented as a Snakemake workflow and is available at: https://gitlab.univ-nantes.fr/bird_pipeline_registry/magneto. IMPORTANCEGenome-resolved metagenomics has led to the discovery of previously untapped biodiversity within the microbial world. As the development of computational methods for the recovery of genomes from metagenomes continues, existing strategies need to be evaluated and compared to eventually lead to standardized computational workflows. In this study, we compared commonly used assembly and binning strategies and assessed their performance using both simulated and real metagenomic datasets. We propose a novel approach to automate co-assembly, avoiding the requirement for a priori knowledge to combine metagenomic information. The comparison against a previous co-assembly approach demonstrates a strong impact of this step on genome binning results, but also the benefits of informing co-assembly for improving the quality of recovered genomes. MAGNETO integrates complementary assembly-binning strategies to optimize genome reconstruction and provides a complete reads-to-genomes workflow for the growing microbiome research community.

bioinformatics↗