bioRxiv Science⌕ Search

Biology subjects

Hoang, T. L.

Publications and source records attributed to Hoang, T. L..

2 recordsLinked to original sources

Effect of dataset partitioning strategies for evaluating out-of-distribution generalisation for predictive models in biochemistry

AO_SCPLOWBSTRACTC_SCPLOWQuantifying model generalization to out-of-distribution data has been a longstanding challenge in machine learning. Addressing this issue is crucial for leveraging machine learning in scientific discovery, where models must generalize to new molecules or materials. Current methods typically split data into train and test sets using various criteria -- temporal, sequence identity, scaffold, or random cross-validation -- before evaluating model performance. However, with so many splitting criteria available, existing approaches offer limited guidance on selecting the most appropriate one, and they do not provide mechanisms for incorporating prior knowledge about the target deployment distribution(s). To tackle this problem, we have developed a novel metric, AU-GOOD, which quantifies expected model performance under conditions of increasing dissimilarity between train and test sets, while also accounting for prior knowledge about the target deployment distribution(s), when available. This metric is broadly applicable to biochemical entities, including proteins, small molecules, nucleic acids, or cells; as long as a relevant similarity function is defined for them. Recognizing the wide range of similarity functions used in biochemistry, we propose criteria to guide the selection of the most appropriate metric for partitioning. We also introduce a new partitioning algorithm that generates more challenging test sets, and we propose statistical methods for comparing models based on AU-GOOD. Finally, we demonstrate the insights that can be gained from this framework by applying it to two different use cases: developing predictors for pharmaceutical properties of small molecules, and using protein language models as embeddings to build biophysical property predictors.

bioinformatics↗

A comparative study of isothermal nucleic acid amplification methods for SARS-CoV-2 detection at point of care

COVID-19, caused by the novel coronavirus SARS-CoV-2, has spread worldwide and put most of the world under lockdown. Despite that there have been emergently approved vaccines for SARS-CoV-2, COVID-19 cases, hospitalizations, and deaths have remained rising. Thus, rapid diagnosis and necessary public health measures are still key parts to contain the pandemic. In this study, the colorimetric isothermal nucleic acid amplification tests (iNAATs) for SARS-CoV-2 detection based on loop-mediated isothermal amplification (LAMP), cross-priming amplification (CPA), and polymerase spiral reaction (PSR) were designed and evaluated. The three methods showed the same limit of detection (LOD) value of 1 copy of the targeted gene per reaction. However, for the direct detection of SARS-CoV-2 genomic-RNA, LAMP outperformed both CPA and PSR, exhibiting the LOD value of roughly 43.14 genome copies/reaction. The results can be read with the naked eye within 45 minutes, without cross-reactivity to closely related coronaviruses. Moreover, the direct detection of SARS-CoV-2 RNA in simulated patient specimens by iNAATs was also successful. Finally, the ready-to-use lyophilized reagents for LAMP reactions were shown to maintain the sensitivity and LOD value of the liquid assays. The results indicate that the colorimetric lyophilized LAMP kit developed herein is highly suitable for detecting SARS-CoV-2 nucleic acids at point-of-care.

molecular biology↗