bioRxiv Science⌕ Search

Biology subjects

Garcia-Seisdedos, D.

Publications and source records attributed to Garcia-Seisdedos, D..

3 recordsLinked to original sources

Integrated view and comparative analysis of baseline protein expression in mouse and rat tissues

The increasingly large amount of proteomics data in the public domain enables, among other applications, the combined analyses of datasets to create comparative protein expression maps covering different organisms and different biological conditions. Here we have reanalysed public proteomics datasets from mouse and rat tissues (14 and 9 datasets, respectively), to assess baseline protein abundance. Overall, the aggregated dataset contained 23 individual datasets, including a total of 211 samples coming from 34 different tissues across 14 organs, comprising 9 mouse and 3 rat strains, respectively. In all cases, we studied the distribution of canonical proteins between the different organs. The number of canonical proteins per dataset ranged from 273 (tendon) and 9,715 (liver) in mouse, and from 101 (tendon) and 6,130 (kidney) in rat. Then, we studied how protein abundances compared across different datasets and organs for both species. As a key point we carried out a comparative analysis of protein expression between mouse, rat and human tissues. We observed a high level of correlation of protein expression among orthologs between all three species in brain, kidney, heart and liver samples, whereas the correlation of protein expression was generally slightly lower between organs within the same species. Protein expression results have been integrated into the resource Expression Atlas for widespread dissemination. Author summaryWe have reanalysed 23 baseline mass spectrometry-based public proteomics datasets stored in the PRIDE database. Overall, the aggregated dataset contained 211 samples, coming from 34 different tissues across 14 organs, comprising 9 mouse and 3 rat strains, respectively. We analysed the distribution of protein expression across organs in both species. We also studied how protein abundances compared across different datasets and organs for both species. Then we performed gene ontology and pathway enrichment analyses to identify enriched biological processes and pathways across organs. We also carried out a comparative analysis of baseline protein expression across mouse, rat and human tissues, observing a high level of expression correlation among orthologs in all three species, in brain, kidney, heart and liver samples. To disseminate these findings, we have integrated the protein expression results into the resource Expression Atlas.

bioinformatics↗

An integrated view of baseline protein expression in human tissues

The availability of proteomics datasets in the public domain, and in the PRIDE database in particular, has increased dramatically in recent years. This unprecedented large-scale availability of data provides an opportunity for combined analyses of datasets to get organism-wide protein abundance data in a consistent manner. We have reanalysed 24 public proteomics datasets from healthy human individuals, to assess baseline protein abundance in 31 organs. We defined tissue as a distinct functional or structural region within an organ. Overall, the aggregated dataset contains 67 healthy tissues, corresponding to 3,119 mass spectrometry runs covering 498 samples, coming from 489 individuals. We compared protein abundances between the different organs and studied the distribution of proteins across organs. We also compared the results with data generated in analogous studies. We also performed gene ontology and pathway enrichment analyses to identify organ-specific enriched biological processes and pathways. As a key point, we have integrated the protein abundance results into the resource Expression Atlas, where it can be accessed and visualised either individually or together with gene expression data coming from transcriptomics datasets. We believe this is a good mechanism to make proteomics data more accessible for life scientists.

bioinformatics↗

Implementing the re-use of public DIA proteomics datasets: from the PRIDE database to Expression Atlas

The number of mass spectrometry (MS)-based proteomics datasets in the public domain keeps increasing, particularly those generated by Data Independent Acquisition (DIA) approaches such as SWATH-MS. Unlike Data Dependent Acquisition datasets, the re-use of DIA datasets has been rather limited to date, despite its high potential, due to the technical challenges involved. We introduce a (re-)analysis pipeline for public SWATH-MS datasets which includes a combination of metadata annotation protocols, automated workflows for MS data analysis, statistical analysis, and the integration of the results into the Expression Atlas resource. Automation is orchestrated with Nextflow, using containerised open analysis software tools, rendering the pipeline readily available and reproducible. To demonstrate its utility, we reanalysed 10 public DIA datasets from the PRIDE database, comprising 1,278 SWATH-MS runs. The robustness of the analysis was evaluated, and the results compared to those obtained in the original publications. The final expression values were integrated into Expression Atlas, making SWATH-MS experiments more widely available and combining them with expression data originating from other proteomics and transcriptomics datasets.

bioinformatics↗