bioRxiv Science⌕ Search

Biology subjects

Shu, K.

Publications and source records attributed to Shu, K..

3 recordsLinked to original sources

Proteogenomics analysis of human tissues using pangenomes

The genomics landscape is evolving with the emergence of pangenomes, challenging the conventional single-reference genome model. The new human pangenome reference provides an extra dimension by incorporating variations observed in different human populations. However, the increasing use of pangenomes in human reference databases poses challenges for proteomics, which currently relies on UniProt canonical/isoform-based reference proteomics. Including more variant information in human proteomes, such as small and long open reading frames and pseudogenes, prompts the development of complex proteogenomics pipelines for analysis and validation. This study explores the advantages of pangenomes, particularly the human reference pangenome, on proteomics, and large-scale proteogenomics studies. We reanalyze two large human tissue datasets using the quantms workflow to identify novel peptides and variant proteins from the pangenome samples. Using three search engines SAGE, COMET, and MSGF+ followed by Percolator we analyzed 91,833,481 MS/MS spectra from more than 30 normal human tissues. We developed a robust deep-learning framework to validate the novel peptides based on DeepLC, MS2PIP and pyspectrumAI. The results yielded 170142 novel peptide spectrum matches, 4991 novel peptide sequences, and 3921 single amino acid variants, corresponding to 2367 genes across five population groups, demonstrating the effectiveness of our proteogenomics approach using the recent pangenome references.

bioinformatics↗

Intron Losses and Gains in Nematodes: Not Eccentric at All

The evolution of spliceosomal introns has been widely studied among various eukaryotic groups. Researchers nearly reached the consensuses on the pattern and the mechanisms of intron losses and gains across eukaryotes. However, according to previous studies that analyzed a few genes or genomes of nematodes, Nematoda seem to be an eccentric group. Taking advantage of the recent accumulation of sequenced genomes, we carried out an extensive analysis on the intron losses and gains using 104 nematodes genomes across all the five Clades of the phylum. Nematodes have a wide range of intron density, from less than one to more than nine per 1kbp coding sequence. The rates of intron losses and gains exhibit significant heterogeneity both across different nematode lineages and across different evolutionary stages of the same lineage. The frequency of intron losses far exceeds that of intron gains. Five pieces of evidence supporting the model of cDNA-mediated intron loss have been observed in ten Caenorhabditis species, the dominance of the precise intron losses, frequent loss of adjacent introns, and high-level expression of the intron-lost genes, preferential losses of short introns, and the preferential losses of introns close to 3'-ends of genes. Like studies in most eukaryotic groups, we cannot find the source sequences for the limited number of intron gains detected in the Caenorhabditis genomes. All the results indicate that nematodes are a typical eukaryotic group rather than an outlier in intron evolution.

genomics↗

Phoenix Enhancer: an online service/tool for proteomics data mining using clustered spectra

MotivationSpectrum clustering has been used to enhance proteomics data analysis: some originally unidentified spectra can potentially be identified and individual peptides can be evaluated to find potential mis-identifications by using clusters of identified spectra. The Phoenix Enhancer provides an infrastructure to analyze tandem mass spectra and the corresponding peptides in the context of previously identified public data. Based on PRIDE Cluster data and a newly developed pipeline, four functionalities are provided: i) evaluate the original peptide identifications in an individual dataset, to find low confidence peptide spectrum matches (PSMs) which could correspond to mis-identifications; ii) provide confidence scores for all originally identified PSMs, to help users evaluate their quality (complementary to getting a global false discovery rate); iii) identify potential new PSMs for originally unidentified spectra; and iv) provide a collection of browsing and visualization tools to analyze and export the results. In addition to the web based service, the code is open-source and easy to re-deploy on local computers using Docker containers. AvailabilityThe service of Phoenix Enhancer is available at http://enhancer.ncpsb.org. All source code is freely available in GitHub (https://github.com/phoenix-cluster/) and can be deployed in the Cloud and HPC architectures. Contactbaimz@cqupt.edu.cn Supplementary informationSupplementary data are available online.

bioinformatics↗