bioRxiv Science⌕ Search

Biology subjects

Gualdi, F.

Publications and source records attributed to Gualdi, F..

2 recordsLinked to original sources

PREDICTING GENE DISEASE ASSOCIATIONS WITH KNOWLEDGE GRAPH EMBEDDINGS FOR DISEASES WITH CURTAILED INFORMATION

Knowledge graph embeddings (KGE) are a powerful technique used in the biological domain to represent biological knowledge in a low dimensional space. However, a deep understanding of these methods is still missing, and in particular the limitations for diseases with reduced information on gene-disease associations. In this contribution, we built a knowledge graph (KG) by integrating heterogeneous biomedical data and generated KGEs by implementing state-of-the-art methods, and two novel algorithms: DLemb and BioKG2Vec. Extensive testing of the embeddings with unsupervised clustering and supervised methods showed that our novel approaches outperform existing algorithms in both scenarios. Our results indicate that data preprocessing and integration influence the quality of the predictions and that the embeddings efficiently encodes biological information when compared to a null model. Finally, we employed KGE to predict genes associated with Intervertebral disc degeneration (IDD) and showed that functions relevant to the disease are enriched in the genes prioritized from the model GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=110 SRC="FIGDIR/small/575314v1_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@1ba98b7org.highwire.dtl.DTLVardef@1801d6borg.highwire.dtl.DTLVardef@b87a7org.highwire.dtl.DTLVardef@f70cf8_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Genopyc: a python library for investigating the genomic basis of complex diseases

MotivationUnderstanding the genetic basis of complex diseases is a paramount challenge in modern genomics. However, current tools often lack the versatility to efficiently analyze the intricate relation-ships between genetic variations and disease outcomes. To address this, we introduce Genopyc, a novel Python library designed for comprehensive investigation of the genetics underlying complex dis-eases. Genopyc offers an extensive suite of functions for heterogeneous data mining and visualization, enabling researchers to delve into and integrate biological information from large-scale genomic da-tasets with ease. ResultsIn this study, we present the Genopyc library through application to real-world genome wide association studies variants. Using Genopyc to investigate variants associated to intervertebral disc degeneration (IDD) enabled a deeper understanding of the potential dysregulated pathways involved in the disease, which can be explored and visualized by exploiting the functionalities featured in the package. Genopyc emerges as a powerful asset for researchers, fostering advancements in the un-derstanding of complex diseases and thus paving the way for more targeted therapeutic interventions. Availability: Genopyc is available at pip (https://pypi.org/project/genopyc/) and the source code of Genopyc is available at https://github.com/freh-g/genopyc Contactfrancesco.gualdi01@estudiant.upf.edu Supplementary informationsupplementary data are available at Bioinformatics online.

bioinformatics↗