bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.08.22.609095

Machine learning model to predict risk assessment of a child inheriting a genetic disorder

Abstract

Advancements in Machine Learning (ML) have revolutionised precision medicine, particularly in predicting and preventing genetic disorders. This study provides a comprehensive analysis of ML models designed for the risk assessment of genetic disorders in children, based on clinical history and pedigree analysis. Leveraging a diverse dataset of familial genetic profiles, clinical outcomes, and medical histories, we employed ML algorithms to identify inheritance patterns, assess genetic risks for Mendelian disorders, and predict disease recurrence. The analysis incorporates several ML algorithms, including Gradient Boosting, XGBoost, Random Forest, Logistic Regression, Naive Bayes, and Support Vector Machines (SVM), with the Gradient Boosting model achieving the highest mean cross-validation score of over 0.99. Designed for primary care physicians and healthcare professionals, this model aids in genetic counselling by predicting genetic disorder recurrence based on family history. The paper also addresses ethical and legal considerations, emphasising the importance of genetic counselling and informed decision-making. This tool is intended to support, not replace, medical professionals. This work advances personalised risk assessment for Mendelian single-gene disorders, including autosomal dominant, autosomal recessive, and X-linked recessive disorders, contributing to the field of genomic medicine and facilitating effective family planning strategies. Author summaryGenetic disorders are posing significant socio-economic, health, and psychological burdens due to the risk of inheritance, which can impact future generations with health complications. Genetic counselling has emerged as an important tool to help manage this issue, offering guidance to susceptible families. Recent advancements in Artificial Intelligence (AI) and Machine Learning (ML) have introduced new avenues for predicting the risk of genetic disorders with greater accuracy. The aim was to leverage machine learning models that would predict the risk assessment for single gene chromosomal Mendelian disorders. To train these models, patient data was collected from a clinical geneticist at a renowned hospital, all of whom had confirmed genetic testing results. This data was used to train the models that predicted the likelihood of the next child inheriting a genetic disorder. The initial results are promising, demonstrating a good fit with the data. However, it should be noted that larger sample sizes are needed to improve the accuracy of any model. With more extensive data, its predictive capabilities can be significantly enhanced. This tool has the potential to be a resource for genetic counsellors and primary healthcare physicians, aiding them in providing more accurate risk assessments and personalised guidance to families.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ramaswamy, M., Senthilnathan, S., Saravanan, A. R. I., M, S., Sivashanmugam, K.. 2024-08-22. Machine learning model to predict risk assessment of a child inheriting a genetic disorder. https://doi.org/10.1101/2024.08.22.609095

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Utilizing single-cell data for per-cell type eQTL mapping in the human pancreas

Aims/hypothesis The human pancreas is a central organ for metabolic regulation that is comprised of diverse cell types that uniquely contribute to its function. Previous studies have performed expression quantitative trail loci (eQTL) discovery in either whole pancreas or in pancreatic islets, but due to differences between pancreatic cell types, this approach does not reveal cell type-specific effects. In this study, we sought to either implicate the cell type of action for known eQTLs or identify new eQTLs that may have been masked in bulk studies by performing eQTL discovery in individual pancreatic cell types. Methods We clustered 153,018 single-cell RNA sequencing (scRNA-seq) data from 71 pancreatic islet donors from the Human Pancreas Analysis Program (HPAP). We performed eQTL discovery in six pancreatic cell types using this resource directly. We further utilized this single cell resource as a reference to deconvolute bulk pancreatic RNA sequencing data from 305 Genotype Tissue Expression (GTEx) project donors and performed eQTL discovery in four pancreatic cell types. Finally, we performed fine-mapping and co-localization of pancreatic cell type eQTLs with metabolic GWAS to connect our findings to metabolic disease risk. Results From analyzing 71 individuals with single cell profiles, we identified 112 unique eGenes across six pancreatic cell types, 99 of which had been identified previously and 13 unique to this study. From the deconvoluted eQTLs, we identified 3,134 unique eGenes across four pancreatic cell types, 116 of which were unique to our study. Fine-mapping and co-localization of eQTLs with metabolic GWAS yielded key leads that warrant further investigation, such as the association of rs2168101 with LMO1 expression in alpha cells. Conclusions/interpretation We identified new signals that were previously not found in bulk pancreatic eQTL studies and potential cell type of action for several signals that were identified previously. Although there are limitations to the power, and therefore, discoverability of this study, it provides insights into how individual pancreatic cells differently contribute to metabolic disease.

genetics↗

MOD-scTWAS: Leveraging gene co-expression for single-cell transcriptome-wide association studies

Transcriptome-wide association studies (TWAS) provide an effective framework for identifying genes associated with complex traits. Population-scale single-cell transcriptomic data enable genetically regulated expression (GReX) prediction and TWAS analyses at cell-type resolution, but the predictive performance of existing single-cell TWAS methods remains limited. Here, we develop MOD-scTWAS, a module-based method that jointly models GReX for genes within co-expression modules to borrow information across genes. Starting from a generative model for single-cell gene expression, MOD-scTWAS accounts for the heteroscedasticity and cross-gene correlation of individual-level pseudobulk expression in joint GReX prediction. In cross-validation analyses of the OneK1K dataset, MOD-scTWAS achieved higher mean GReX prediction accuracy than scTWAS across all 14 cell types and increased the number of imputable genes. When applied to TWAS analyses of UK Biobank quantitative hematological traits, MOD-scTWAS identified more significant cell type-gene-trait associations than scTWAS. These results demonstrate the potential of leveraging gene co-expression through joint modeling to improve cell-type-specific GReX prediction and TWAS discovery.

genetics↗

Generation of a transgenic cephalopod

Coleoid cephalopods (cuttlefish, octopus, and squid) are marine mollusks with elaborate nervous systems that support a diverse repertoire of complex behaviors. These include the neural control of the color, pattern, and texture of the skin, facilitating both adaptive camouflage and innate patterning that may reflect internal state. The development of transgenic cephalopods expressing fluorescent proteins, optogenetic actuators, and reporters of neural activity would contribute a new and important technology to cephalopod biology. The generation of transgenic cephalopods, however, has remained a major challenge. Here, we report the development of stable transgenic dwarf cuttlefish (Ascarosepion bandense) expressing ubiquitous nuclear-localized mScarlet, a red fluorescent protein. We evaluated multiple strategies for transgenesis, and established cuttlefish lines using both CRISPR and the transposons Sleeping Beauty and Minos. The stable expression of transgenes enabled live imaging of cell dynamics during embryonic development. The Minos transposon emerged as the most efficient transgenesis strategy and is adaptable to promoters and transgenes of choice. These strategies now enable the generation of diverse genetic tools for mechanistic studies of cephalopod biology.

genetics↗