bioRxiv ScienceSearch

Biology subjects

Garcia-Ruiz, S.

Publications and source records attributed to Garcia-Ruiz, S..

4 recordsLinked to original sources

Detection of pathogenic splicing events from RNA-sequencing data using dasper

Although next-generation sequencing technologies have accelerated the discovery of novel gene-to-disease associations, many patients with suspected Mendelian diseases still leave the clinic without a genetic diagnosis. An estimated one third of these patients will have disorders caused by mutations impacting splicing. RNA-sequencing has been shown to be a promising diagnostic tool, however few methods have been developed to integrate RNA-sequencing data into the diagnostic pipeline. Here, we introduce dasper, an R/Bioconductor package that improves upon existing tools for detecting aberrant splicing by using machine learning to incorporate disruptions in exon-exon junction counts as well as coverage. dasper is designed for diagnostics, providing a rank-based report of how aberrant each splicing event looks, as well as including visualization functionality to facilitate interpretation. We validate dasper using 16 patient-derived fibroblast cell lines harbouring pathogenic variants known to impact splicing. We find that dasper is able to detect pathogenic splicing events with greater accuracy than existing LeafCutterMD or z-score approaches. Furthermore, by only applying a broad OMIM gene filter (without any variant-level filters), dasper is able to detect pathogenic splicing events within the top 10 most aberrant identified for each patient. Since using publicly available control data minimises costs associated with incorporating RNA-sequencing into diagnostic pipelines, we also investigate the use of 504 GTEx fibroblast samples as controls. We find that dasper leverages publicly available data effectively, ranking pathogenic splicing events in the top 25. Thus, we believe dasper can increase diagnostic yield for a pathogenic splicing variants and enable the efficient implementation of RNA-sequencing for diagnostics in clinical laboratories.

bioinformatics

Leveraging omic features with F3UTER enables identification of unannotated 3'UTRs for synaptic genes

There is growing evidence for the importance of 3 untranslated region (3UTR) dependent regulatory processes. However, our current human 3UTR catalogue is incomplete. Here, we developed a machine learning-based framework, leveraging both genomic and tissue-specific transcriptomic features to predict previously unannotated 3UTRs. We identify unannotated 3UTRs associated with 1,513 genes across 39 human tissues, with the greatest abundance found in brain. These unannotated 3UTRs were significantly enriched for RNA binding protein (RBP) motifs and exhibited high human lineage-specificity. We found that brain-specific unannotated 3UTRs were enriched for the binding motifs of important neuronal RBPs such as TARDBP and RBFOX1, and their associated genes were involved in synaptic function and brain- related disorders. Our data is shared through an online resource F3UTER (https://astx.shinyapps.io/F3UTER/). Overall, our data improves 3UTR annotation and provides novel insights into the mRNA-RBP interactome in the human brain, with implications for our understanding of neurological and neurodevelopmental diseases.

bioinformatics

Human brain mitochondrial-nuclear cross-talk is cell-type specific and is perturbed by neurodegeneration

Mitochondrial dysfunction contributes to the pathogenesis of many neurodegenerative diseases as mitochondria are essential to neuronal function. The mitochondrial genome encodes a small number of core respiratory chain proteins, whereas the vast majority of mitochondrial proteins are encoded by the nuclear genome. Here we focus on establishing a profile of nuclear-mitochondrial transcriptional relationships in healthy human central nervous system tissue data, before examining perturbations of these processes in Alzheimer&#8217s disease using transcriptomic data originating from affected human brain tissue. Through cross-central nervous system analysis of mitochondrial-nuclear gene pair relationships, we find that the cell type composition underlies regional variation, and variation is driven at the subcellular level by heterogeneity of nuclear-mitochondrial coordination in post-synaptic regions. We show that nuclear genes causally implicated in sporadic Parkinson&#8217s disease and Alzheimer&#8217s disease show much stronger relationships with the mitochondrial genome than expected by chance, and that nuclear-mitochondrial relationships are significantly perturbed in Alzheimer&#8217s disease cases, particularly amongst genes involved in synaptic and lysosomal pathways. Finally, we present MitoNuclearCOEXPlorer, a web tool designed to allow users to interrogate and visualise key mitochondrial-nuclear relationships in multi-dimensional brain data. We conclude that mitochondrial-nuclear relationships differ significantly across regions of the healthy brain, which appears to be driven by the functional specialisation of different cell types. We also find that mitochondrial-nuclear co-expression in critical pathways is disrupted in Alzheimer&#8217s disease, potentially implicating the regulation of energy balance and removal of dysfunctional mitochondria in the etiology or progression of the disease and making the case for the relevance of bi-genomic co-ordination in the pathogenesis of neurodegenerative diseases.

genomics

Modeling multifunctionality of genes with secondary gene co-expression networks in human brain provides novel disease insights

MotivationCo-expression networks are a powerful gene expression analysis method to study how genes co-express together in clusters with functional coherence that usually resemble specific cell type behaviour for the genes involved. They can be applied to bulk-tissue gene expression profiling and assign function, and usually cell type specificity, to a high percentage of the gene pool used to construct the network. One of the limitations of this method is that each gene is predicted to play a role in a specific set of coherent functions in a single cell type (i.e. at most we get a single for each gene). We present here GMSCA (Gene Multifunctionality Secondary Co-expression Analysis), a software tool that exploits the co-expression paradigm to increase the number of functions and cell types ascribed to a gene in bulk-tissue co-expression networks. ResultsWe applied GMSCA to 27 co-expression networks derived from bulk-tissue gene expression profiling of a variety of brain tissues. Neurons and glial cells (microglia, astrocytes and oligodendrocytes) were considered the main cell types. Applying this approach, we increase the overall number of predicted triplets by 46.73%. Moreover, GMSCA predicts that the SNCA gene, traditionally associated to work mainly in neurons, also plays a relevant function in oligodendrocytes. AvailabilityThe tool is available at GitHub,https://github.com/drlaguna/GMSCA as open source software. ImplementationGSMCA is implemented in R.

bioinformatics