bioRxiv Science⌕ Search

Biology subjects

Corbett, M.

Publications and source records attributed to Corbett, M..

2 recordsLinked to original sources

Knowledge Inclusive Machine Learning for Disease Gene Prioritisation

The predictive performance of machine learning models depends on the context available to them. In disease gene prioritisation, this context comprises two forms: specific context from sample-level experimental data, such as gene expression and protein-protein interaction networks, and general context from accumulated and curated biological knowledge capturing established relationships among genes, diseases, and pathways. Neither is sufficient alone: experimental data are sensitive to dataset-specific noise and lack broader biological grounding, while curated knowledge lacks the resolution required for gene-level discrimination. Consequently, most machine learning approaches relying solely on experimental data risk learning spurious correlations rather than underlying biology. Here we introduce Knowledge Inclusive Machine Learning (KIML), a paradigm that integrates both context types within a unified analytical pipeline. KIML combines experimental data with two types of general context: literature-derived representations from PubMed and structured biomedical knowledge graphs. We evaluate the approach on Developmental and Epileptic Encephalopathy and benchmark it against recent methods using publicly available datasets. Performance is assessed using temporal-split evaluation and biological evaluations, including ontology enrichment analysis. KIML consistently outperforms existing approaches, providing improved predictive accuracy and biologically meaningful insights. Furthermore, the framework generates interpretable explanations of gene prioritisation and demonstrates strong generalisability across six additional diseases.

genetics↗

Optical genome mapping enables accurate repeat expansion testing

Short tandem repeats (STRs) are amongst the most abundant class of variations in human genomes and are meiotically and mitotically unstable which leads to expansions and contractions. STR expansions are frequently associated with genetic disorders, with the size of expansions often correlating with the severity and age of onset. Therefore, being able to accurately detect the total repeat expansion length and to identify potential somatic repeat instability is important. Current standard of care (SOC) diagnostic assays include laborious repeat-primed PCR-based tests as well as Southern blotting, which are unable to precisely determine long repeat expansions and/or require a separate set-up for each locus. Sequencing-based assays have proven their potential for the genome-wide detection of repeat expansions but have not yet replaced these diagnostic assays due to their inaccuracy to detect long repeat expansions (short-read sequencing) and their costs (long-read sequencing). Here, we tested whether optical genome mapping (OGM) can efficiently and accurately identify the STR length and assess the stability of known repeat expansions. We performed OGM for 85 samples with known clinically relevant repeat expansions in DMPK, CNBP and RFC1, causing myotonic dystrophy type 1 and 2 and cerebellar ataxia, neuropathy and vestibular areflexia syndrome (CANVAS), respectively. After performing OGM, we applied three different repeat expansion detection workflows, i.e. manual de novo assembly, local guided assembly (local-GA) and molecule distance script of which the latter two were developed as part of this study. The first two workflows estimated the repeat size for each of the two alleles, while the third workflow was used to detect potential somatic instability. The estimated repeat sizes were compared to the repeat sizes reported after the SOC and concordance between the results was determined. All except one known repeat expansions above the pathogenic repeat size threshold were detected by OGM, and allelic differences were distinguishable, either between wildtype and expanded alleles, or two expanded alleles for recessive cases. An apparent strength of OGM over current SOC methods was the more accurate length measurement, especially for very long repeat expansion alleles, with no upper size limit. In addition, OGM enabled the detection of somatic repeat instability, which was detected in 9/30 DMPK, 23/25 CNBP and 4/30 RFC1 samples, leveraging the analysis of intact, native DNA molecules. In conclusion, for tandem repeat expansions larger than [~]300 bp, OGM provides an efficient method to identify exact repeat lengths and somatic repeat instability with high confidence across multiple loci simultaneously, enabling the potential to provide a significantly improved and generic genome-wide assay for repeat expansion disorders.

genomics↗