bioRxiv Science⌕ Search

Biology subjects

Kohli, A.

Publications and source records attributed to Kohli, A..

7 recordsLinked to original sources

Integrated histopathologic modeling of detailed tumor subtypes and actionable biomarkers

Accurate cancer subtyping with accompanying molecular characterization is critical for precision oncology. While machine learning approaches have been applied to both digital pathology and cancer genomics, previous work has been limited in sample size and has typically aggregated granular cancer subtypes into coarse groupings, likely obfuscating informative molecular and prognostic associations and phenotypic variation of more detailed tumor subtypes. Accordingly, we collated 378,123 hematoxylin and eosin (H&E)-stained whole-slide images (WSIs) with matched targeted DNA clinical sequencing results and OncoTree detailed cancer subtypes from a real-world cohort of 71,142 patients. Using this scaled, granular dataset and a cancer subtype knowledge graph, we developed Mosaic: a family of calibrated machine learning models using H&E WSI embeddings to classify tumors and identify molecular phenotypes across 163 detailed subtypes. The cancer subtyping module (Aeon) achieved an area under the receiver operating characteristic curve (AUROC) of 0.992 overall, with 161/163 subtypes reaching an AUROC [≥] 0.90 and improved performance over a state-of-the-art genomics-based classifier. The genomic inference module (Paladin) achieved an AUROC [≥] 0.80 for 167 pairs of detailed subtypes and genomic targets. We further used the learned histopathologic representations to i) identify key associations of the histopathologic embeddings with clinical biomarkers; ii) identify unsupervised sub-clusters of tumors with genomic determinants of tumor phenotype; iii) specify granular diagnoses for cancers of unknown primary, evaluated by genomic associations and expected clinical outcome distributions; iv) annotate functional significance for variants of uncertain significance (VUS); and v) identify cases that mimic the phenotypic effect of known DNA variants on H&E in the absence of detectable DNA alterations. Taken together, this work advances our understanding of phenotypic variation of granular tumor subtypes, their relevance to enhanced diagnostics, and their potential utility in risk stratification with multimodal machine learning in cancer.

cancer biology↗

Chromatix: a differentiable, GPU-accelerated wave-optics library

Modern microscopy methods incorporate computational modeling as an integral part of the imaging process, either to solve inverse problems or optimize the optical system design itself. These methods often depend on differentiable optics simulations, yet no standardized framework exists--forcing computational optics researchers to repeatedly and independently implement simulations with limited reusability and performance. These common problems limit the potential impact of computational optics as a field. Here we present Chromatix: an open-source, GPU-accelerated, differentiable wave optics simulation library. Chromatix builds on JAX to democratize fast, parallelized simulation of diverse optical systems and expand the design space in computational optics. Chromatix standardizes a growing collection of optical elements and propagation methods allowing a broad range of applications, which we demonstrate here for snapshot microscopy, holography, and phase retrieval. We demonstrate speed improvements of 2-6x on a single GPU and up to 22x on 8 GPUs.

bioinformatics↗

Future flooding tolerant rice germplasm: resilience afforded beyond Sub1A gene

Developing high-yielding, flood-tolerant rice varieties is essential for enhancing productivity and livelihoods in flood-prone ecologies. We explored genetic avenues beyond the well-known SUB1A gene to improve flood resilience in rice. We screened a collection of 6,274 elite genotypes from IRRIs germplasm repository for submergence and stagnant flooding tolerance over multiple seasons and years. This rigorous screening identified 89 outstanding elite genotypes, among which thirty-seven exhibited high submergence tolerance, surpassing the survival rate of SUB1A introgression genotypes by 40-50%. Thirty-five genotypes showed significant tolerance to stagnant flooding, and 17 demonstrated dual tolerance capabilities, highlighting their adaptability to varying flood conditions. The genotypes identified have a broader genetic diversity and harbor 86 key QTLs and genes related to traits such as grain quality, grain yield, herbicide resistance, and various biotic and abiotic traits, highlighting the richness of the identified elite collection. Besides germplasm, we introduce an innovative breeding approach called Transition from Trait to Environment (TTE). TTE leverages a parental pool of high-performing genotypes with complete submergence tolerance to drive population improvement and enable genomic selection in the flood breeding program. Our approach of TTE achieved a remarkable 65% increase in genetic gain for submergence tolerance, with the resulting fixed breeding genotypes demonstrating exceptional performance in flood-prone environments of India and Bangladesh. The elite genotypes identified herein represent invaluable genetic resources for the global rice research community. By adopting the TTE approach, which is trait agonistic, we establish a robust framework for developing more resilient genotypes using advanced breeding tools. Plain Language SummaryTo address climate challenges, an urgent focus is necessary to identify and develop flood-tolerant rice varieties, particularly for flood-prone ecosystems across Asia and Africa. We screened 6,274 elite genotypes from IRRIs germplasm and identified 89 promising lines with improved tolerance to submergence and stagnant flooding. Among these, 37 demonstrated 40-50% greater submergence tolerance than SUB1A introgression lines, 35 exhibited stagnant flooding tolerance, and 17 showed dual tolerance. These genotypes contain 86 key QTLs and genes associated with yield, grain quality, and biotic and abiotic tolerance traits. A new breeding strategy, the Transition from Trait to Environment (TTE) approach, was developed. We achieved a genetic gain of 65% for submergence tolerance in rice using this method. The newly identified germplasm provides invaluable genetic resources for the global rice research community to develop flood-tolerant rice genotypes. Core ideas The SUB1A gene, enabling rice to survive underwater for 14 days, marked a significant breakthrough. We have identified elite genotypes with submergence tolerance significantly surpassing the SUB1A gene-mediated tolerance. The diverse elite genotypes identified harbor 86 key genes and QTLs that affect various traits positively. Developed a unique breeding strategy for implementing population improvement in challenging environments. The new breeding strategy demonstrated a genetic gain of 65% for submergence tolerance.

plant biology↗

Hybrid Transformer and Neural NetworkConfiguration for Protein Classification UsingAmino Acids

This study introduces a hybrid machine learning model for classifying proteins, developed to address the complexities of protein sequence and structural analysis. Utilizing an architecture that combines a lightweight transformer with a concurrent neural network, the hybrid model leverages both sequential and intrinsic physical properties of proteins. Trained on a comprehensive dataset from the Research Collaboratory for Structural Bioinformatics Protein Data Bank, the model demonstrates a classification accuracy of 95%, outperforming existing methods by at least 15%. The high accuracy achieved demonstrates the potential of this approach to innovate protein classification, facilitating advancements in drug discovery and the development of personalized medicine. By enabling precise protein function prediction, the hybrid model allows for specialized strategies in therapeutic targeting and the exploration of protein dynamics in biological systems. Future work will focus on enhancing the models generalizability across diverse datasets and exploring the integration of more machine learning techniques to refine predictive capabilities further. The implications of this research offer potential breakthroughs in biomedical research and the broader field of protein engineering.

bioengineering↗

Potential Opportunities of Modeling Bioavailability for Monoclonal Antibodies: An Overview of mAbs and the current challenges of mAb development

With a growing market size, and a large variety of applications, monoclonal antibody technology adoption and clinical usage is at an all-time high. This review article seeks to explore 10 monoclonal antibodies (mAbs) and their mechanism of action, specifically their pharmacodynamic (PD) and pharmacokinetic (PK) properties, and use a machine learning model with various parameters to assess whether the mAb has adequate bioavailability when delivered subcutaneously. This is an investigation of drug optimization and patient outcomes when transitioning from traditional IV administrations to subcutaneous injections. The machine learning model is an extension based on a paper by Han Lou and Michael Hageman, Machine Learning Attempts for Predicting Human Subcutaneous Bioavailability of Monoclonal Antibodies, where they took 10 mAbs and analyzed 45 different features. To further extend this paper, we took an additional 10 monoclonal antibodies that were delivered subcutaneously, and took into account their dosage concentration as an extension to traditional PK properties. By including additional mAbs and dosage, a more sophisticated model can be produced with high scalability to deep learning modalities.

bioengineering↗

Multi-omics of a rice population identifies genes and genomic regions in rice that bestow low glycemic index and high protein content

To address the growing incidences of increased diabetes and to meet the daily protein requirements, we developed low glycemic index (GI) rice varieties with protein yield exceeding 14%. In the development of recombinant inbred lines using Samba Mahsuri and IR36 amylose extender as parental lines, we identified quantitative trait loci (QTLs) and genes associated with low GI, high amylose content (AC), and high protein content (PC). By integrating genetic techniques with classification models, this comprehensive approach identified candidate genes on chromosome 2 (qGI2.1/qAC2.1 spanning the region from 18.62Mb to 19.95Mb), exerting influence on low GI and high amylose. Notably, the phenotypic variant with high value was associated with the recessive allele of the starch branching enzyme 2b (sbeIIb). The genome-edited sbeIIb line confirmed low GI phenotype in milled rice grains. Further, combinations of alleles from the highly significant SNPs from the targeted associations and epistatically interacting genes showed ultra-low GI phenotypes with high amylose and high protein. Metabolomics analysis of rice with varying AC, PC, and GI revealed that the superior lines of high AC and PC, and low GI were preferentially enriched in glycolytic and amino acid metabolism, whereas the inferior lines of low AC and PC and high GI were enriched with fatty acid metabolism. The high amylose high protein RIL (HAHP_101) was enriched in essential amino acids like lysine. Such lines may be highly relevant for food product development to address diabetes and malnutrition. Significance StatementThe increasing global incidence of diabetes calls for the development of diabetic friendly healthier rice. In this study, we developed recombinant inbred rice lines with milled rice exhibiting ultra-low to low glycemic index and high protein content from the cross between Samba Mahsuri and IR36 amylose extender. We performed comprehensive genomics and metabolomics complemented with modeling analyses emphasizing the importance of OsSbeIIb along with additional candidate genes whose variations allowed us to produce target rice lines with lower glycemic index and high protein content in a high-yielding background. These lines represent an important breeding resource to address food and nutritional security.

systems biology↗

DeepMap: A deep learning-based model with four-line code for prediction-based breeding in crops

Prediction of phenotype through genotyping data using the emerging machine or deep learning technology has been proven successful in genomic prediction. We present here a graphical processing unit (GPU) enabled DeepMap configurable deep learning-based python package for the genomic prediction of quantitative phenotype traits. We found that deep learning captures non-linear patterns more efficiently than conventional statistical methods. Furthermore, we suggest an additional module inclusion of epistasis interactions and training of the model on Graphical Processing Units (GPUs) in addition to Central Processing Unit (CPU) to enhance efficiency and increase the models performance. We developed and demonstrated the application of DeepMap using a 3K rice genome panel and 1K-Rice Custom Amplicon (1kRiCA) data for several phenotypic traits including days to 50% flowering (DTF), number of productive tillers (NPT), panicle length (PL), plant height (PH), and plot yield (PY). We have found that DeepMap outperformed the best existing state-of-the-art models by giving higher predictive correlation and low mean squared error for the datasets studied. This prediction performance was higher than other compared models in the range of 13-31%. Similarly for Dataset-2, significantly higher predictions were observed than the compared models (16-20% higher prediction ability). On Dataset-3, we have also shown the better and versatile performance of our model across crops (wheat, maize, and soybean) for yield and yield-related traits. This demonstrates the potentiality of the framework and ease of use for future research in crop improvement. The DeepMap is accessible at https://test.pypi.org/project/DeepMap-1.0/. Short SummaryDeepMap is a deep learning-based breeder-friendly python package to perform genomic prediction. It utilizes epistatic interactions for data augmentation and outperforms the existing state-of-the-art machine/deep learning models such as Bayesian LASSO, GBLUP, DeepGS, and dualCNN. DeepMap developed for rice and tested across crops such as maize, wheat, soybean etc.

genomics↗