bioRxiv Science⌕ Search

Biology subjects

Neekhra, B.

Publications and source records attributed to Neekhra, B..

3 recordsLinked to original sources

Robust Prediction of Patient-Specific Cancer Hallmarks Using Neural Multi-Task Learning: a model development and validation study

BackgroundAccurate quantification of cancer hallmark activity is essential for understanding tumor progression, tailoring treatments, and improving patient outcomes. Traditional methods, such as histopathological grading and immunohistochemistry for protein expression, often overlook the complex interplay between cancer cells and the tumor microenvironment and provide limited insight into hallmark-specific mechanisms. We aimed to develop OncoMark, a high-throughput deep learning-enabled neural multi-task learning framework capable of systematically quantifying integrative hallmarks activities using transcriptomics data from routine tumor biopsies. MethodsIn this study, we acquired single-cell transcriptomics data from 941 tumor samples across 14 tissue types, comprising nearly 3.1 million cells from 56 studies conducted worldwide, to form a large multicenter dataset. Our model employs a supervised neural multi-task learning method designed to predict multiple cancer hallmarks present in the biopsy samples simultaneously. The OncoMark model was developed and tested on 90% of the studies (patients from 51 studies) using repeated five-fold cross-validation performed twice. For further evaluation, the model was assessed on the remaining 10% of the studies (patients from 5 studies) that were excluded from the initial training and testing dataset. Additionally, we included patients from publicly available datasets, including TCGA, GTEx, ANTE, MET500, POG570, CCLE, TARGET, and PCAWG to validate the models performance. The primary objective was to evaluate the performance of the model in identifying cancer hallmarks in cancer datasets and ensure no hallmark predictions were made in normal samples across the four prespecified groups: (i) internal test set, (ii) external test set, (iii) normal samples (real-world), and (iv) cancer samples (real-world). FindingsOncoMark demonstrated exceptional performance in predicting cancer hallmark states, achieving near-perfect accuracy across internal test data and five external test datasets. Internal testing consistently showed accuracy, precision, recall, and F1 scores exceeding 99%, underscoring the models reliability across hallmarks. External test further confirmed these findings, with accuracy, precision, recall, F1 scores, and balanced accuracy consistently exceeding 96{middle dot}6%, and multiple datasets achieving perfect scores, highlighting the models exceptional generalizability and robustness. Specificity tests using GTEx and ANTE datasets accurately classified normal tissues, while sensitivity analysis on TCGA, MET500, CCLE, TARGET, PCAWG, and POG570 datasets effectively identified cancer hallmarks. InterpretationWe developed an AI-based framework that enables accurate, efficient, and cost-effective quantification of cancer hallmark activity directly from transcriptomics data. The framework demonstrated significant potential as an assistive tool for guiding personalized treatment strategies and advancing the clinical management of cancer patients. FundingAshoka University, S.N. Bose National Centre for Basic Sciences, Mphasis F1 Foundation, DST SERB Core Research Grant. Research in ContextO_ST_ABSEvidence before this studyC_ST_ABSWe conducted an extensive literature search using Google Scholar and PubMed without language restrictions, employing search terms such as "(Predicting OR Classifying OR Annotating) and (cancer hallmarks) AND (Deep OR Machine Learning) OR (Artificial Intelligence OR AI)." While there have been advancements in molecular oncology and computational methodologies over the two decades since the concept of cancer hallmarks was first introduced, a comprehensive machine learning or deep learning framework to annotate all cancer hallmarks simultaneously from tumor biopsy samples remains to be developed. Additionally, the scarcity of hallmark-annotated datasets has posed a significant challenge, hindering the development of robust predictive models. Added value of this studyThis study introduces OncoMark, a novel high-throughput neural multi-task learning (N-MTL) framework designed to predict all cancer hallmark activities simultaneously from biopsy samples. OncoMark addresses the lack of annotated hallmark-specific data by generating synthetic biopsy (pseudo-bulk) datasets annotated with hallmark activity, meticulously modeled to reflect real-world tumor biology while maintaining clinical relevance. The framework employs a multi-task learning approach to capture interdependencies among hallmarks, advancing beyond isolated predictions to offer a holistic view of tumor biology. Validation on five independent datasets comprising 95 patient samples demonstrated its generalizability and reproducibility. Further external validation using eight datasets, encompassing over 11,679 cancer and 8348 normal patient samples, reinforced its robustness. To promote clinical integration, a user-friendly web-based tool was developed, enabling seamless access for oncologists and researchers. Implications of all the available evidenceThe OncoMark framework represents a transformative advancement in cancer diagnostics and treatment planning. By enabling accurate and reproducible prediction of all hallmark activities simultaneously from biopsy samples, this model paves the way for precision oncology at scale. Its ability to systematically capture hallmark interdependencies provides deeper insights into tumor behavior, guiding the development of individualized targeted therapies. The incorporation of a web-based interface ensures the accessibility of this innovation to clinicians worldwide, bridging the gap between computational oncology and clinical practice. Following further validation and integration into healthcare workflows, OncoMark has the potential to improve cancer outcomes by delivering timely, cost-effective, and precise tumor analyses, facilitating informed therapeutic decision-making with unparalleled precision. Cancer progression is driven by a set of well-defined biological principles--collectively termed the "hallmarks of cancer"--yet current diagnostic approaches seldom incorporate these distinct molecular features into clinical practice. Despite substantial progress in molecular oncology, traditional methods like histopathological grading and immunohistochemical assays often fail to capture the complex interplay between cancer cells and the tumor microenvironment, emphasizing the need for robust computational frameworks capable of systematically quantifying hallmark-specific activity. Here, we address this gap by developing OncoMark, a high-throughput neural multi-task learning (N-MTL) framework designed to simultaneously quantify hallmark activities in tumor biopsies using transcriptomics data. We show that OncoMark achieves near-perfect accuracy, precision, recall, and F1 scores (>99%) in cross-validation, with external validation consistently exceeding 96.6% on five independent datasets. Further evaluation on eight additional datasets--including large-scale cancer cohorts (TCGA, MET500, CCLE, TARGET, PCAWG, POG570) and normal tissue datasets (GTEx, ANTE)--demonstrated high specificity for normal samples and robust sensitivity for hallmark prediction in cancer. By delivering a comprehensive and cost-effective molecular portrait of tumor biology and providing a user-friendly web platform accessible at https://oncomark-ai.hf.space/, OncoMark has the potential to guide tailored treatment strategies and advance precision oncology. More broadly, this framework signifies a transformative step toward routine hallmark-based diagnostics, promising to improve patient outcomes by facilitating timely and precise tumor profiling.

cancer biology↗

Comprehensive Enumeration of Cancer Stem-like Cell Heterogeneity Using Deep Neural Network

Cancer stem cells (CSCs), a distinct subpopulation within tumors, are pivotal in driving treatment resistance and tumor recurrence, posing substantial challenges to conventional therapeutic strategies. Precise quantification and profiling of these cells are essential for improving cancer treatment outcomes. We present ACSCeND, an advanced deep neural network model accompanied by a robust workflow, specifically developed to quantify cellular compositions from bulk RNA-seq data, enabling accurate CSC profiling. By integrating bulk RNA-seq data with insights derived from single-cell RNA-seq datasets, ACSCeND effectively captures the diversity and hierarchical organization of tumor-resident cell states, alongside cell-specific gene expression profiles (GEPs). Compared to current tissue deconvolution models, ACSCeND exhibits superior performance, achieving significantly higher Concordance Correlation Coefficient (CCC) values and lower Root Mean Square Error (RMSE) across various pseudobulk and real-world bulk tissue samples. Application of ACSCeND to TCGA and PRECOG datasets reveals a strong association between CSC abundance and poorer disease-free survival outcomes, underscoring the clinical relevance of CSCs in cancer progression. Furthermore, cell-specific GEPs for distinct CSC states unveil novel molecular signatures and illuminate the origins of CSC-driven tumor heterogeneity. In summary, ACSCeND provides a powerful, scalable platform for high-throughput quantification of cellular compositions and distinct potency states within normal tissues as well as highly heterogeneous tissues, such as tumors.

genomics↗

Next-Gen Profiling of Tumor-resident Stem Cells using Machine learning

Tumor-resident stem cells, also known as cancer stem cells (CSCs), constitute a subgroup within tumors, play a crucial role in fostering resistance to treatment and the recurrence of tumors, and pose significant challenges for conventional therapeutic methods. Existing approaches for identifying CSCs face notable hurdles related to scalability, reproducibility, and technical consistency across different cancer types due to the adaptable nature of CSCs. In this context, we introduce OSCORP, an innovative machine-learning-driven approach. This methodology quantifies and identifies CSCs, achieving almost 99% accuracy using biopsy bulk RNAseq data. OSCORP leverages genetic similarities between normal and cancer stem cells. By categorizing CSCs into four distinct yet dynamic potency states, this approach provides insights into the differentiation landscape of CSCs, unveiling previously undisclosed facets of tumor heterogeneity. In evaluations conducted on patient samples across 22 cancer types, OSCORP revealed clinical, transcriptomic, and immunological signatures associated with each CSC state. It has emerged as a comprehensive tool for understanding and addressing the complexities of cancer stem cells. Ultimately, OSCORP opens up new possibilities for more effective personalized cancer therapies and holds the potential to serve as a clinical tool for monitoring patient-specific CSC changes during treatment or follow-up care.

cancer biology↗