bioRxiv Science⌕ Search

bioRxiv · 10.1101/2024.03.28.587184

Establishing the foundations for a data-centric AI approach for virtual drug screening through a systematic assessment of the properties of chemical data

Abstract

Researchers have adopted model-centric artificial intelligence (AI) approaches in cheminformatics by using newer, more sophisticated AI methods to take advantage of growing chemical libraries. It has been shown that complex deep learning methods outperform conventional machine learning (ML) methods in QSAR and ligand-based virtual screening1-3 but such approaches generally lack explanability. Hence, instead of developing more sophisticated AI methods (i.e., pursuing a model-centric approach), we wanted to explore the potential of a data-centric AI paradigm for virtual screening. A data-centric AI is an intelligent system that would automatically identify the right type of data to collect, clean and curate for later use by a predictive AI and this is required given the large volumes of chemical data that exist in chemical databases - PubChem alone has over 100 million unique compounds. However, a systematic assessment of the attributes and properties of suitable data is needed. We show here that it is not the result of deficiencies in current AI algorithms but rather, poor understanding and erroneous use of chemical data that ultimately leads to poor predictive performance. Using a new benchmark dataset of BRAF ligands that we developed, we show that our best performing predictive model can achieve an unprecedented accuracy of 99% with a conventional ML algorithm (SVM) using a merged molecular representation (Extended + ECFP6 fingerprints), far surpassing past performances of virtual screening platforms using sophisticated deep learning methods. Thus, we demonstrate that it is not necessary to resort to the use of sophisticated deep learning algorithms for virtual screening because conventional ML can perform exceptionally well if given the right data and representation. We also show that the common use of decoys for training leads to high false positive rates and its use for testing will result in an over-optimistic estimation of a models predictive performance. Another common practice in virtual screening is defining compounds that are above a certain pharmacological threshold as inactives. Here, we show that the use of these so-called inactive compounds lowers a models sensitivity/recall. Considering that some target proteins have a limited number of known ligands, we wanted to also observe how the size and composition of the training data impact predictive performance. We found that an imbalance training dataset where inactives outnumber actives led to a decrease in recall but an increase in precision, regardless of the model or molecular representation used; and overall, we observed a decrease in the models accuracy. We highlight in this study some of the considerations that one needs to take into account in future development of data-centric AI for CADD.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chong, A., Phua, S.-X., Xiao, Y., Ng, W. Y., Li, H. Y., Goh, W. W. B.. 2024-03-31. Establishing the foundations for a data-centric AI approach for virtual drug screening through a systematic assessment of the properties of chemical data. https://doi.org/10.1101/2024.03.28.587184

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Ex vivo human tumor slices more accurately predict patient responses to an oncolytic virus than in vivo mouse models

Immunotherapies, including oncolytic viruses (OV), are promising therapies that can enhance anti-tumor immune responses. However, preclinical success of immunotherapies in mouse models has not always translated to clinical benefit in cancer patients. This study compared preclinical efficacy and mechanism of action for ASP9801, a vaccinia virus expressing IL-7 and IL-12, using mouse models of colorectal cancer (CRC) in vivo and in human organotypic tumor slice models ex vivo. The murine surrogate for ASP9801 significantly reduced tumor volumes in treated and abscopal tumors in two different CRC models in vivo (MC38 and RO100). Treatment efficacy was accentuated when combined with anti-PD1 treatment, and single-cell RNA sequencing analysis revealed depletion of tumor cells and increased T cell infiltration and activation in both treated and abscopal tumors. However, human tissue analysis ex vivo (E-slices) using PDX models and patient samples showed that ASP9801 is not effective in CRC, consistent with clinical trial results. On the other hand, ASP9801 was highly effective in GBM, indicating indication-specific efficacy of ASP9801, and how E-slice assays can be used to identify treatment-sensitive indications. This study demonstrates the superiority of E-slices over mouse models for predicting clinical response and its utility in planning clinical trials.

cancer biology↗

Immune-cell depleted diffuse large B-cell lymphomas have reduced expression of MHC class I

Immunotherapy has transformed treatment for many cancers. In the aggressive and genetically heterogeneous diffuse large B-cell lymphoma (DLBCL), CD19 CAR T-cell therapy is highly effective, whereas immune checkpoint blockade has shown limited benefit. Loss of MHC expression is a common mechanism to escape T-cell cytotoxicity, and loss of MHC class I (MHC-I) and II are frequent in DLBCL. We applied imaging mass cytometry to diagnostic biopsies from younger, high-risk DLBCL patients to map the tumor microenvironment (TME) spatial architecture in relation to tumor cell MHC expression, mutational status, transcriptomic and proteomic profiles. Neighborhood analyses identified four TME subtypes: immune-cell depleted and three immune-infiltrated types (mixed, CD4 T cell-rich, CD8 T-cell/macrophage-rich). Depleted cases had shorter overall survival (p = 0.033) and increased expression of proteins involved in DNA replication and proliferation markers compared to infiltrated cases. Tumor cell MHC-I expression was heterogeneous. Cases with low frequency of MHC-I-pos tumor cells were enriched for the depleted TME type. MHC-I-pos tumor cells were surrounded by CD4 and CD8 T cells and M1 macrophages, whereas MHC-I-neg tumor cells were closer to other MHC-I-neg tumor cells. These findings suggest that TME-based classification incorporating tumor cell MHC-I status may improve individualized immunotherapy selection.

cancer biology↗

Cross-species analysis links cell-cell communication rewiring to NOTCH2 during serous endometrial carcinogenesis

Cell-cell interactions shape the fate of mutant cells during cancer initiation but how these interactions evolve during progression to pathologically recognizable lesions remain poorly understood. Here, we investigated cell-cell communication during serous endometrial carcinoma (SEC; also known as uterine serous carcinoma) development using a lineage-traceable mouse model and cross-species analyses of the mouse and human neoplastic endometrium. In mice, the early, pre-dysplastic stage was marked by a global decrease in inferred cell-cell interactions, followed by extensive communication network rewiring during neoplastic progression. Pathway-specific analysis revealed a similar pattern for NOTCH signaling, with NOTCH2 emerging as the dominant NOTCH receptor in Trp53/Rb1-mutant immature epithelial cells. Functionally, NOTCH2 promoted the outgrowth of more proliferative mutant organoids. Cross-species transcriptomic analysis identified conserved immature epithelial states in mouse and human neoplastic endometrial epithelium. In human tissues, NOTCH2 was overexpressed in serous endometrial intraepithelial carcinoma, a precursor of SEC, and in overt SEC. Furthermore, elevated NOTCH2 expression was associated with poor patient survival. These findings link cell-cell communication rewiring during experimental SEC development to conserved neoplastic epithelial states and identify NOTCH2 as an early marker and a potential target of disease interception.

cancer biology↗