bioRxiv Science⌕ Search

Biology subjects

De Paoli, F.

Publications and source records attributed to De Paoli, F..

2 recordsLinked to original sources

PhenoXtract: combining Large Language Model and Knowledge Graph embedding to extract phenotypes from clinical descriptions

MotivationStandardized phenotypic descriptions are essential for accurate diagnosis, yet clinicians and researchers face challenges in manually extracting and mapping phenotypes from scientific literature or patient clinical records to the Human Phenotype Ontology. Recent advances in deep learning offer new opportunities for automation. We developed PhenoXtract, a novel phenotype extraction approach that combines Large Language Models and Knowledge Graph embedding. PhenoXtract is a multistep pipeline that takes clinical descriptions as input, extracts candidate phenotype entities using large language models, and maps them to terms from an enriched version of the Human Phenotype Ontology, processed as a knowledge graph. ResultsEvaluation against expert-curated ground-truth datasets show a recall of 0.70 and precision of 0.85 for PhenoXtract, demonstrating concordance with manually extracted phenotypes, with a computation time of 10-20 seconds for each text analyzed. Moreover, PhenoXtract surpasses rule-based and deep learning-based state-of-the-art tools in two out of the three ground-truth datasets evaluated. These results suggest that hybrid approaches combining Large Language Models and Knowledge Graph embeddings represent a promising direction for automated clinical phenotyping at scale. Contactsberardelli@engenome.com

genomics↗

Digenic variant interpretation with hypothesis-driven explainable AI

MotivationThe digenic inheritance hypothesis holds the potential to enhance diagnostic yield in rare diseases. Computational approaches capable of accurately interpreting and prioritizing digenic combinations based on the probands phenotypic profiles and familial information can provide valuable assistance to clinicians during the diagnostic process. ResultsWe have developed diVas, a hypothesis-driven machine learning approach that can effectively interpret genomic variants across different gene pairs. DiVas demonstrates strong performance both in classifying and prioritizing causative pairs, consistently placing them within the top positions across 11 real cases (achieving 73% sensitivity and a median ranking of 3). Additionally, diVas exploits Explainable Artificial Intelligence (XAI) to dissect the digenic disease mechanism for predicted positive pairs. Availability and ImplementationPrediction results of the diVas method on a high-confidence, comprehensive, manually curated dataset of known digenic combinations are available at oliver.engenome.com.

bioinformatics↗