bioRxiv Science⌕ Search

Biology subjects

Dash, T.

Publications and source records attributed to Dash, T..

7 recordsLinked to original sources

Non-canonical regulation of the plasma membrane copper transporter CTR1 through modulation of membrane mechanical properties

We describe a non-canonical, membrane receptor-like regulation of the human copper transporter CTR1 in response to copper stimuli. CTR1 is the sole high-affinity trimeric plasma-membrane copper-importing channel that self-regulates by undergoing endocytosis to limit copper uptake. We observed that preceding copper-induced endocytosis, CTR1 forms clusters on the plasma membrane, a phenomenon that is typically observed in membrane receptors. We deciphered the mechanism of CTR1 clustering and studied its ramifications on plasma membrane physical properties harboring the clusters that favors endocytosis. Membrane tension and fluctuation are fundamental regulators of pre- and post-endocytic events. Using coarse-grain MD simulation and coupled Interference Reflection Microscopy-Total Internal Reflection Fluorescence Microscopy we demonstrated that CTR1 clusters induce positive membrane curvature, increase in local membrane tension and decrease in local membrane fluctuation; alterations that favors formation of endocytic pits. Clustering is facilitated by copper-sequestering Methionine rich extracellular amino-terminus of CTR1. MD-simulations and IRM-TIRF imaging revealed that CTR1 clustering is facilitated by membrane cholesterol, depletion of which delays CTR1 endocytosis. CTR1 clustering promotes clathrin-coated pit formation that engages recruitment of adapter protein AP-2. To summarize, we report a hitherto unknown pre-endocytic receptor-like phenomenon of ligand-induced clustering of a metal channel, that in-turn regulates self-endocytosis by modulating membrane properties.

cell biology↗

Identifying a logical specification and a program for an LLM-based generator of lead molecules

Our interest is in the generation of "lead" molecules in early-stage drug design. Leads are small molecules (ligands) that can bind to a part of pre-specified target and also satisfy multiple physico-chemical constraints. We propose using techniques developed in Inductive Logic Programming (ILP) to identify a logical specification of feasible molecules; and then using this specification to construct a program that uses a large language model (LLM) to generate new molecules. We ensure the program constructed is correct, in the sense that every molecule generated by the program is feasible according the specification. Our focus is on contributing to on-going drug-discovery research on novel inhibitors for Dopamine {beta}-hydroxylase (DBH), an enzyme that plays a pivotal role in several diseases related to the brain and the heart. We find molecules comparable in affinity to the latest generation drugs currently in clinical trials, and chemical assessment of synthesisablity of the molecules generated. For completeness, we also provide results obtained on the classic benchmark datasets used in recent work reported in [1].

bioinformatics↗

Predicting gene expression using millions of yeastpromoters reveals cis-regulatory logic

MotivationGene regulation involves complex interactions between multiple transcription factors. While early attempts to train deep neural networks to predict gene expression were limited to naturally occurring promoter sequences, the advent of gigantic parallel reporter assays has expanded available training data by orders of magnitude. Despite these advances, a clear understanding of how to use deep learning to study gene regulation is still lacking. MethodHere we investigate the complex association between gene promoters and expression in S. cerevisiae using Camformer, a residual convolutional neural network that ranked 4th in the Random Promoter DREAM Challenge 2022. We present the original model trained on 6.7 million random promoter sequences and investigate 270 alternative models to determine what factors contribute most to model performance. Finally, we use explainable AI to uncover regulatory signals. ResultsWe show that Camformer accurately decodes the association between promoters and gene expression (r2 = 0.914 {+/-} 0.003,{rho} = 0.962 {+/-} 0.002) and provides a substantial improvement over previous state of the art. Furthermore, we show that a much smaller model with approximately 90% fewer parameters than the original model can achieve a high predictive performance. Using Grad-CAM and in silico mutagenesis, we demonstrate that the model learns both individual motifs and their hierarchy. For example, while an IME1 motif on its own increases gene expression, the co-occurrence of a UME6 motif provides a switch to strongly reduce gene expression. Thus, deep learning models such as Camformer can provide detailed insights into cis-regulatory logic. Availability and ImplementationThe data and code used and developed in our experiments are publicly available at: https://github.com/Bornelov-lab/Camformer.

bioinformatics↗

A comprehensive multi-omics study reveals potential prognostic and diagnostic biomarkers for colorectal cancer

BackgroundColorectal cancer (CRC) is a complex disease with diverse genetic alterations and causes 10% of cancer-related deaths worldwide. Understanding its molecular mechanisms is essential for identifying potential biomarkers and therapeutic targets for its effective management. MethodWe integrated copy number alterations (CNA) and mutation data via their differentially expressed genes termed as candidate genes (CGs) computed using bioinformatics approaches. Then, using the CGs, we perform Weighted correlation network analysis (WGCNA) and utilise several hazard models such as Univariate Cox, Least Absolute Shrinkage and Selection Operator (LASSO) Cox and multivariate Cox to identify the key genes involved in CRC progression. We used different machine-learning models to demonstrate the discriminative power of selected hub genes among normal and CRC (early and late-stage) samples. ResultsThe integration of CNA with mRNA expression identified over 3000 CGs, including CRC-specific driver genes like MYC and APC. In addition, pathway analysis revealed that the CGs are mainly enriched in endocytosis, cell cycle, wnt signalling and mTOR signalling pathways. Hazard models identified four key genes, CASP2, HCN4, LRRC69 and SRD5A1, that were significantly associated with CRC progression and predicted the 1-year, 3-years, and 5-years survival times. WGCNA identified seven hub genes: DSCC1, ETV4, KIAA1549, NOP56, RRS1, TEAD4 and ANKRD13B, which exhibited strong predictive performance in distinguishing normal from CRC (early and late-stage) samples. ConclusionsIntegrating regulatory information with gene expression improved early versus latestage prediction. The identified potential prognostic and diagnostic biomarkers in this study may guide us in developing effective therapeutic strategies for CRC management.

bioinformatics↗

Generating Novel Leads for Drug Discovery using LLMs with Logical Feedback

Large Language Models (LLMs) can be used as repositories of biological and chemical information to generate pharmacological lead compounds. However, for LLMs to focus on specific drug targets typically require experimentation with progressively more refined prompts. Results thus become dependent not just on what is known about the target, but also on what is known about the prompt-engineering. In this paper, we separate the prompt into domain-constraints that can be written in a standard logical form, and a simple text-based query. We investigate whether LLMs can be guided, not by refining prompts manually, but by refining the the logical component automatically, keeping the query unchanged. We describe an iterative procedure LMLF ("Language Models with Logical Feedback") in which the constraints are progressively refined using a logical notion of generalisation. On any iteration, newly generated instances are verified against the constraint, providing "logical-feedback" for the next iterations refinement of the constraints. We evaluate LMLF using two well-known targets (inhibition of the Janus Kinase 2; and Dopamine Receptor D2); and two different LLMs (GPT-3 and PaLM). We show that LMLF, starting with the same logical constraints and query text, can guide both LLMs to generate potential leads. We find: (a) Binding affinities of LMLF-generated molecules are skewed towards higher binding affinities than those from existing baselines; LMLF results in generating molecules that are skewed towards higher binding affinities than without logical feedback; (c) Assessment by a computational chemist suggests that LMLF generated compounds may be novel inhibitors. These findings suggest that LLMs with logical feedback may provide a mechanism for generating new leads without requiring the domain-specialist to acquire sophisticated skills in prompt-engineering.

bioinformatics↗

An AI-assisted Investigation of Tumor-Associated Macrophages and their Polarization in Colorectal Cancer

Tumor-associated Macrophages (or TAMs) are amongst the most common cells that play a significant role in the initiation and progression of colorectal cancer (CRC). [Ghosh et al., 2023] have built a Boolean-logic dependent model to propose a set of gene signatures capable of identifying macrophage polarization states. The signature, called the Signature of Macrophage Reactivity and Tolerance (SMaRT), comprises of 338 human genes (equivalently, 298 mouse genes). The SMaRT signature was constructed using datasets that were not specialized towards any particular disease. To specifically investigate macrophage polarization in CRC, in this paper, we (a) perform a comprehensive analysis of the SMaRT signature on single-cell human and mouse colorectal cancer RNA-seq datasets and (b) adopt transfer learning to construct a "refined" SMaRT signature that specifically characterizes TAM polarization in the CRC tumor microenvironment. Towards validation of our refined gene signature, we use: (a) 5 RNA-seq datasets derived from single-cell human datasets; and (b) 5 large-cohort microarray datasets from humans. Furthermore, we propose the translational potential of our refined gene signature while investigating microsatellite stability and CpG island methylator phenotype (CIMP) in colorectal cancer. Overall, our refined gene signature and its extensive validation provide a path for its adoption in clinical practice in diagnosing colorectal cancer and associated attributes. Availability and ImplementationThe data, codes, and software packages used in our research are linked and shared publicly at https://github.com/tirtharajdash/TAMs-CRC.

bioinformatics↗

Using Domain-Knowledge to Assist Lead Discovery in Early-Stage Drug Design

We are interested in generating new small molecules which could act as inhibitors of a biological target, when there is limited prior information on target-specific inhibitors. This form of drug-design is assuming increasing importance with the advent of new disease threats for which known chemicals only provide limited information about target inhibition. In this paper, we propose the combined use of deep neural networks and Inductive Logic Programming (ILP) that allows the use of symbolic domain-knowledge (B) to explore the large space of possible molecules. Assuming molecules and their activities to be instances of random variables X and Y, the problem is to draw instances from the conditional distribution of X, given Y, B (DX|Y,B). We decompose this into the constituent parts of obtaining the distributions DX|B and DY|X,B, and describe the design and implementation of models to approximate the distributions. The design consists of generators (to approximate DX|B and DX|Y,B) and a discriminator (to approximate DY|X,B). We investigate our approach using the well-studied problem of inhibitors for the Janus kinase (JAK) class of proteins. We assume first that if no data on inhibitors are available for a target protein (JAK2), but a small numbers of inhibitors are known for homologous proteins (JAK1, JAK3 and TYK2). We show that the inclusion of relational domain-knowledge results in a potentially more effective generator of inhibitors than simple random sampling from the space of molecules or a generator without access to symbolic relations. The results suggest a way of combining symbolic domain-knowledge and deep generative models to constrain the exploration of the chemical space of molecules, when there is limited information on target-inhibitors. We also show how samples from the conditional generator can be used to identify potentially novel target inhibitors.

molecular biology↗