bioRxiv Science⌕ Search

Biology subjects

Kuznetsova, K.

Publications and source records attributed to Kuznetsova, K..

5 recordsLinked to original sources

Retention time and fragmentation predictors increase confidence in variant peptide identification

Precision medicine focuses on adapting care to the individual profile of patients, e.g. accounting for their unique genetic makeup. Being able to account for the effect of genetic variation on the proteome holds great promises towards this goal. However, identifying the protein products of genetic variation using mass spectrometry has proven very challenging. Here we show that the identification of variant peptides can be improved by the integration of retention time and fragmentation predictors into a unified proteogenomic pipeline. By combining these intrinsic peptide characteristics using the search-engine post-processor Percolator, we demonstrate improved discrimination power between correct and incorrect peptide-spectrum matches. Our results demonstrate that the drop in performance that is induced when expanding a protein sequence database can be compensated, and hence enabling efficient identification of genetic variation products in proteomics data. We anticipate that this enhancement of proteogenomic pipelines can provide a more refined picture of the unique proteome of patients, and thereby contribute to improving patient care.

bioinformatics↗

A systematic mapping of the genomic and proteomic variation associated with monogenic diabetes

AimsMonogenic diabetes is characterized as a group of diseases caused by rare variants in single genes. Multiple genes have been described to be responsible for monogenic diabetes, but the information on the variants is not unified among different resources. In this work, we aimed to develop an automated pipeline that collects all the genetic variants associated with monogenic diabetes from different resources, unify the data and translate the genetic sequences to the proteins. MethodsThe pipeline developed in this work is written in Python with the use of Jupyter notebook. It consists of 6 modules that can be implemented separately. The translation step is performed using the ProVar tool also written in Python. All the code along with the intermediate and final results is available for public access and reuse. ResultsThe resulting database had 2701 genomic variants in total and was divided into two levels: the variants reported to have an association with monogenic diabetes and the variants that have evidence of pathogenicity. Of them, 2565 variants were found in the ClinVar database and the rest 136 were found in the literature showing that the overlap between resources is not absolute. ConclusionsWe have developed an automated pipeline for collecting and harmonizing data on genetic variants associated with monogenic diabetes. Furthermore, we have translated variant genetic sequences into protein sequences accounting for all protein isoforms and their variants. This allows researchers to consolidate information on variant genes and proteins associated with monogenic diabetes and facilitates their study using proteomics or structural biology. Our open and flexible implementation using Jupyter notebooks enables tailoring and modifying the pipeline and its application to other rare diseases. Research in contextO_LIMonogenic diabetes is a group of Mendelian diseases with an autosomal-dominant pattern of inheritance. C_LIO_LIMonogenic diabetes is mainly caused by rare genetic variants that are usually evaluated manually. C_LIO_LIThe data on the variants are stored in several resources and are not unified in terms of the genomic coordinates, alleles, and variant annotation. C_LIO_LIWhat can be done for the systematic evaluation of the variants and their protein consequences? C_LIO_LIIn this work, we have created an automated Jupyter notebook-based pipeline for the collection and unification of the variants associated with monogenic diabetes. C_LIO_LIThe database of the genetic variants was created and translated to all possible variant protein sequences. C_LIO_LIThese results will be used for the analysis of proteomics data and protein structure modeling. C_LI

bioinformatics↗

Targeted disruption of transcription bodies causes widespread activation of transcription

The localization of transcriptional activity in specialized transcription bodies is a hallmark of gene expression in eukaryotic cells. It remains unclear, however, if and how they affect gene expression. Here, we disrupted the formation of two prominent endogenous transcription bodies that mark the onset of zygotic transcription in zebrafish embryos and analysed the effect on gene expression using enriched SLAM-Seq and live-cell imaging. We find that the disruption of transcription bodies results in downregulation of hundreds of genes, providing experimental support for a model in which transcription bodies increase the efficiency of transcription. We also find that a significant number of genes are upregulated, counter to the suggested stimulatory effect of transcription bodies. These upregulated genes have accessible chromatin and are poised to be transcribed in the presence of the two transcription bodies, but they do not go into elongation. Live-cell imaging shows that the disruption of the two large transcription bodies enables these poised genes to be transcribed in ectopic transcription bodies, suggesting that the large transcription bodies sequester a pause release factor. Supporting this hypothesis, we find that CDK9, the kinase that releases paused polymerase II, is highly enriched in the two large transcription bodies. Importantly, overexpression of CDK9 in wild type embryos results in the formation of ectopic transcription bodies and thus phenocopies the removal of the two large transcription bodies. Taken together, our results show that transcription bodies regulate transcription genome-wide: the accumulation of transcriptional machinery creates a favourable environment for transcription locally, while depriving genes elsewhere in the nucleus from the same machinery.

cell biology↗

Nanog organizes transcription bodies

The localization of transcriptional activity in specialized transcription bodies is a hallmark of gene expression in eukaryotic cells. How proteins of the transcriptional machinery come together to form such bodies, however, is unclear. Here, we take advantage of two large, isolated, and long-lived transcription bodies that reproducibly form during early zebrafish embryogenesis, to characterize the dynamics of transcription body formation. Once formed, these transcription bodies are enriched for initiating and elongating RNA polymerase II, as well as the transcription factors Nanog and Sox19b. Analyzing the events leading up to transcription, we find that Nanog and Sox19b cluster prior to transcription, and independently of RNA accumulation. The clustering of transcription factors is sequential; Nanog clusters first, and this is required for the clustering of Sox19b and the initiation of transcription. Mutant analysis revealed that both the DNA-binding domain, as well as one of the two intrinsically disordered regions of Nanog are required to organize the two bodies of transcriptional activity. Taken together, our data suggests that the clustering of transcription factors dictates the formation of transcription bodies. HIGHLIGHTSO_LITranscription factors cluster prior to, and independently of transcription C_LIO_LINanog organizes transcription bodies: it is required for the clustering of Sox19b as well as RNA polymerase II C_LIO_LIThis organizing activity requires its DNA binding domain as well as one of its intrinsically disordered regions C_LIO_LITranscription elongation results in the disassembly of transcription factor clusters C_LI

cell biology↗

Validating amino acid variants in proteogenomics using sequence coverage by multiple reads

Mass spectrometry-based proteome analysis usually implies matching mass spectra of proteolytic peptides to amino acid sequences predicted from nucleic acid sequences. At the same time, due to the stochastic nature of the method when it comes to proteome-wide analysis, in which only a fraction of peptides are selected for sequencing, the completeness of protein sequence identification is undermined. Likewise, the reliability of peptide variant identification in proteogenomic studies is suffering. We propose a way to interpret shotgun proteomics results, specifically in data-dependent acquisition mode, as protein sequence coverage by multiple reads, just as it is done in the field of nucleic acid sequencing for the calling of single nucleotide variants. Multiple reads for each position in a sequence could be provided by overlapping distinct peptides, thus, confirming the presence of certain amino acid residues in the overlapping stretch with much lower false discovery rate than conventional 1%. The source of overlapping distinct peptides are, first, miscleaved tryptic peptides in combination with their properly cleaved counterparts, and, second, peptides generated by several proteases with different specificities after the same specimen is subject to parallel digestion and analyzed separately. We illustrate this approach using publicly available multiprotease proteomic datasets and our own data generated for HEK-293 cell line digests obtained using trypsin, LysC and GluC proteases. From 5000 to 8000 protein groups are identified for each digest corresponding to up to 30% of the whole proteome coverage. Most of this coverage was provided by a single read, while up to 7% of the observed protein sequences were covered two-fold and more. The proteogenomic analysis of HEK-293 cell line revealed 36 peptide variants associated with SNP, seven of which were supported by multiple reads. The efficiency of the multiple reads approach depends strongly on the depth of proteome analysis, the digesting features such as the level of miscleavages, and will increase with the number of different proteases used in parallel proteome digestion. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=75 SRC="FIGDIR/small/475497v1_ufig1.gif" ALT="Figure 1"> View larger version (14K): org.highwire.dtl.DTLVardef@1d6ee2org.highwire.dtl.DTLVardef@5ae8baorg.highwire.dtl.DTLVardef@652216org.highwire.dtl.DTLVardef@1a0d49b_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗