bioRxiv Science⌕ Search

Biology subjects

Halgamuge, S. K.

Publications and source records attributed to Halgamuge, S. K..

2 recordsLinked to original sources

Knowledge Inclusive Machine Learning for Disease Gene Prioritisation

The predictive performance of machine learning models depends on the context available to them. In disease gene prioritisation, this context comprises two forms: specific context from sample-level experimental data, such as gene expression and protein-protein interaction networks, and general context from accumulated and curated biological knowledge capturing established relationships among genes, diseases, and pathways. Neither is sufficient alone: experimental data are sensitive to dataset-specific noise and lack broader biological grounding, while curated knowledge lacks the resolution required for gene-level discrimination. Consequently, most machine learning approaches relying solely on experimental data risk learning spurious correlations rather than underlying biology. Here we introduce Knowledge Inclusive Machine Learning (KIML), a paradigm that integrates both context types within a unified analytical pipeline. KIML combines experimental data with two types of general context: literature-derived representations from PubMed and structured biomedical knowledge graphs. We evaluate the approach on Developmental and Epileptic Encephalopathy and benchmark it against recent methods using publicly available datasets. Performance is assessed using temporal-split evaluation and biological evaluations, including ontology enrichment analysis. KIML consistently outperforms existing approaches, providing improved predictive accuracy and biologically meaningful insights. Furthermore, the framework generates interpretable explanations of gene prioritisation and demonstrates strong generalisability across six additional diseases.

genetics↗

CoPR: Collective Pattern Recognition-a Framework for Microbial Community Activity Analysis

BackgroundMicrobial community activities provide essential information on understanding bacterial communities. Unfortunately, they are generally not directly observable. We rely on longitudinal abundance profiles to get insight into microbial community activities. Often datasets do not have sufficient longitudinal sampling points to successfully apply our algorithms. Hence, in this paper, we are interested in analysing multiple datasets from similar environments to alleviate the aforementioned problem. Furthermore, we wish to see whether collective pattern recognition would enhance our understanding of microbial community activities. ResultsIn this paper, we present CoPR, a framework for collective microbial longitudinal abundance data. Our visualisation shows that a single pattern for temporal abundance variation does not exist. However, it also indicates that even complete individuality does not exist. Consequently, our visualisation highlights the individuality and conformity in the temporal variation of abundance profiles of similar host environments. We also identify different characteristics in the TVAP (Temporal Variation of Abundance Profile) patterns with regards to cohesion and separation. ConclusionsCoPR helps gain essential insights into the microbial communities and their heterogeneity through visualisation tools. This paper also highlights the choice between individuality and conformity in microbial community data analysis.

bioinformatics↗