bioRxiv Science⌕ Search

Biology subjects

Leppert, T.

Publications and source records attributed to Leppert, T..

4 recordsLinked to original sources

The Zea mays PeptideAtlas; a new maize community resource

We developed the Maize PeptideAtlas resource (www.peptideatlas.org/builds/maize) to help solve questions about the maize proteome. Publicly available raw tandem mass spectrometry (MS/MS) data for maize were collected from ProteomeXchange and reanalyzed through a uniform processing and metadata annotation pipeline. These data are from a wide range of genetic backgrounds, including the inbred lines B73 and W22, many hybrids and their respective parents. Samples were collected from field trials, controlled environmental conditions, a range of (a)biotic conditions and different tissues, cell types and subcellular fractions. The protein search space included different maize genome annotations for the B73 inbred line from MaizeGDB, UniProtKB, NCBI RefSeq and for the W22 inbred line. 445 million MS/MS spectra were searched, of which 120 million were matched to 0.37 million distinct peptides. Peptides were matched to 66.2% of the proteins (one isoform per protein coding gene) in the most recent B73 nuclear genome annotation (v5). Furthermore, most conserved plastid- and mitochondrial-encoded proteins (NCBI RefSeq annotations) were identified. Peptides and proteins identified in the other searched B73 genome annotations will aid to improve maize genome annotation. We also illustrate high confidence detection of unique W22 proteins. N-terminal acetylation, phosphorylation, ubiquitination, and three lysine acylations (K-acetyl, K-malonyl, K-hydroxyisobutyryl) were identified and can be inspected through a PTM viewer in PeptideAtlas. All matched MS/MS-derived peptide data are linked to spectral, technical and biological metadata. This new PeptideAtlas is integrated with community resources including MaizeGDB at https://www.maizegdb.org/ and a peptide track in JBrowse. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=119 SRC="FIGDIR/small/572651v2_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@1ec9005org.highwire.dtl.DTLVardef@1e35f34org.highwire.dtl.DTLVardef@7f5b2corg.highwire.dtl.DTLVardef@13acefd_HPS_FORMAT_FIGEXP M_FIG C_FIG

plant biology↗

Detection and editing of the updated plastid- and mitochondrial-encoded proteomes for Arabidopsis with PeptideAtlas

Arabidopsis thaliana Col-0 has plastid and mitochondrial genomes encoding for over one hundred proteins and several ORFs. Public databases (e.g. Araport11) have redundancy and discrepancies in gene identifiers for these organelle-encoded proteins. RNA editing results in changes to specific amino acid residues or creation of start and stop codons for many of these proteins, but the impact of such RNA editing at the protein level is largely unexplored due to the complexities of detection. This study first assembled the non-redundant set of identifiers, their correct protein sequences, and 452 predicted non-synonymous editing sites of which 56 are edited at lower frequency. Accumulation of edited and/or unedited proteoforms was then determined by searching [~]259 million raw MSMS spectra from ProteomeXchange as part of Arabidopsis PeptideAtlas (www.peptideatlas.org/builds/arabidopsis/). All mitochondrial proteins and all except three plastid-encoded proteins (NDHG/NDH6, PSBM, RPS16), but none of the ORFs, were identified; we suggest that all ORFs and RPS16 are pseudogenes. Detection frequencies for each edit site and type of edit (e.g. S to L/F) were determined at the protein level, cross-referenced against the metadata (e.g. tissue), and evaluated for technical challenges of detection.167 predicted edit sites were detected at the proteome level. Minor frequency sites were indeed also edited at low frequency at the protein level. However, except for sites RPL5-22 and CCB382-124, proteins only accumulate in edited form (>98 -100% edited) even if RNA editing levels are well below 100%. This study establishes that RNA editing for major editing sites is required for stable protein accumulation.

plant biology↗

Mapping the Arabidopsis thaliana proteome in PeptideAtlas and the nature of the unobserved (dark) proteome; strategies towards a complete proteome

This study describes a new release of the Arabidopsis thaliana PeptideAtlas proteomics resource providing protein sequence coverage, matched mass spectrometry (MS) spectra, selected PTMs, and metadata. 70 million MS/MS spectra were matched to the Araport11 annotation, identifying [~]0.6 million unique peptides and 18267 proteins at the highest confidence level and 3396 lower confidence proteins, together representing 78.6% of the predicted proteome. Additional identified proteins not predicted in Araport11 should be considered for building the next Arabidopsis genome annotation. This release identified 5198 phosphorylated proteins, 668 ubiquitinated proteins, 3050 N-terminally acetylated proteins and 864 lysine-acetylated proteins and mapped their PTM sites. MS support was lacking for 21.4% (5896 proteins) of the predicted Araport11 proteome - the dark proteome. This dark proteome is highly enriched for certain (e.g. CLE, CEP, IDA, PSY) but not other (e.g. THIONIN, CAP,) signaling peptides families, E3 ligases, TFs, and other proteins with unfavorable physicochemical properties. A machine learning model trained on RNA expression data and protein properties predicts the probability for proteins to be detected. The model aids in discovery of proteins with short-half life (e.g. SIG1,3 and ERF-VII TFs) and completing the proteome. PeptideAtlas is linked to TAIR, JBrowse, PPDB, SUBA, UniProtKB and Plant PTM Viewer.

plant biology↗

The Arabidopsis thaliana PeptideAtlas; harnessing world-wide proteomics data for a comprehensive community proteomics resource

We developed a new resource, the Arabidopsis PeptideAtlas (www.peptideatlas.org/builds/arabidopsis/), to solve central questions about the Arabidopsis proteome, such as the significance of protein splice forms, post-translational modifications (PTMs), or simply obtain reliable information about specific proteins. PeptideAtlas is based on published mass spectrometry (MS) analyses collected through ProteomeXchange and reanalyzed through a uniform processing and metadata annotation pipeline. All matched MS-derived peptide data are linked to spectral, technical and biological metadata. Nearly 40 million out of [~]143 million MSMS spectra were matched to the reference genome Araport11, identifying [~]0.5 million unique peptides and 17858 uniquely identified proteins (only isoform per gene) at the highest confidence level (FDR 0.0004; 2 non-nested peptides [≥] 9 aa each), assigned canonical proteins, and 3543 lower confidence proteins. Physicochemical protein properties were evaluated for targeted identification of unobserved proteins. Additional proteins and isoforms currently not in Araport11 were identified, generated from pseudogenes, alternative start, stops and/or splice variants and sORFs; these features should be considered for updates to the Arabidopsis genome. Phosphorylation can be inspected through a sophisticated PTM viewer. This new PeptideAtlas is integrated with community resources including TAIR, tracks in JBrowse, PPDB and UniProtKB. Subsequent PeptideAtlas builds will incorporate millions more MS data. One sentence summaryA new web resource providing the global community with mass spectrometry-based Arabidopsis proteome information and its spectral, technical and biological metadata integrated with TAIR and JBrowse

plant biology↗