bioRxiv Science⌕ Search

Biology subjects

Shehu, A.

Publications and source records attributed to Shehu, A..

9 recordsLinked to original sources

How Much Does Protein Structure Really Help? A Case Study in Mutation-Induced Stability Prediction

Multimodal neural networks integrating protein language models (PLMs) with structure-derived features are increasingly common for predicting mutation effects, yet fundamental mechanisms remain poorly characterized. In this paper, we articulate and address two key questions: (i) do these architectures exploit mutation-conditioned structural changes, and (ii) does structure provide additive value over PLM-learned sequence embeddings afterall? Using {Delta}Tm (melting temperature shift) prediction as a controlled testbed, a phenotype expected to depend strongly on three-dimensional geometry, we introduce generalizable diagnostic methodologies: systematic channel ablations quantifying each modalitys marginal contribution, and context-radius probing, a novel technique restricting inputs to progressively larger neighborhoods around mutations to spatially localize predictive signal. Across ten independent runs per condition, we find PLM embeddings dominate: removing structure causes minimal performance change, while removing PLMs causes performance collapse. Context-radius probing reveals signal is highly localized; mutation-site-only models recover full-context performance. Critically, comparing wild-type-shared versus mutation-conditioned structural regimes reveals no systematic gain from geometric perturbations, demonstrating that current representations function as static fold priors because downstream featurization attenuates mutation-induced changes. However, structure helps selectively: benefits concentrate in variants with atypical PLM embeddings occupying phenotypically incoherent neighborhoods where sequence-derived priors are locally unreliable. Though focused on a controlled testbed, this work surfaces a key challenge for any protein prediction task where domain knowledge suggests structure should matter: not whether to "add structure," but how to represent and integrate geometry so that it contributes distinct signal beyond strong sequence priors. We provide architecture-agnostic diagnostics to test and quantify when and how explicit structure delivers that added value.

bioinformatics↗

Which pLM to choose?

AO_SCPLOWBSTRACTC_SCPLOWProtein-language models (pLMs) provide a novel means for mapping the protein space. Which of these new maps best advances specific biological analyses, however, is not obvious. To elucidate the principles of model selection, we benchmarked fourteen pLMs, spanning several orders of magnitude in parameter count, across a hundred million protein pairs, to assess how well they capture sequence, structure, and function similarity. For each model, we distinguish inherent information, i.e. signal recoverable from raw-embedding distances, and extractable information, i.e. signal revealed through additional supervised training. Three key results emerge. First, pLM protein representation space is inherently different from the space of biological protein representations, i.e. sequences or structures. Here, a size-performance paradox is salient - mid-scale foundation models are as good as much larger ones in reflecting all tested biological properties. Second, pLM representations compress and store biological information in proportion to model size. That is, a lightweight feed-forward network can be trained on embedding pairs to predict said biological properties well - a capacity dividend. Finally, we observe that a task-specific learning radically reshapes the embedding space, gaining inherent understanding of the task, but garbling any further extractions. In other words, smaller pLMs can provide efficient and compute-light general insight. Larger models are advantageous only when fine-tuning is planned to accomplish a specific task. Furthermore, representations generated by "specialist" models are not immediately generalizable throughout protein biology. Thus, for pLMs, bigger isnt always better.

bioinformatics↗

Rod photoreceptors control the ON vs OFF polarity of cone-signaling neurons

A fundamental feature of the visual system is its ability to detect image contrast. The contrast processing starts in the first synapse of the retina where parallel pathways are established to compute contrast to bright (ON pathway) and dark (OFF pathway) objects, separately transferred to morphologically identified ON and OFF cells throughout the visual system. Here, we found that response polarity in ON and OFF neurons is not fixed but rather switches dynamically to the opposite sign. The switch was not observed in rod-knockout mice, indicating that rods generate the polarity switch. We determined that neither horizontal cells nor rod-signaling pathways were responsible for the switch. Instead, we discovered that EAAT5 glutamate transporters located at photoreceptor terminals were required to produce the polarity switch. Our findings provide a new perspective on the adaptive properties of neural networks and their ability to encode contrast across the visual dynamic range.

neuroscience↗

Efficient High-Throughput DNA Breathing Features Generation Using Jax-EPBD

DNA breathing dynamics--transient base-pair opening and closing due to thermal fluctuations--are vital for processes like transcription, replication, and repair. Traditional models, such as the Extended Peyrard-Bishop-Dauxois (EPBD), provide insights into these dynamics but are computationally limited for long sequences. We present JAX-EPBD, a high-throughput Langevin molecular dynamics framework leveraging JAX for GPU-accelerated simulations, achieving up to 30x speedup and superior scalability compared to the original C-based EPBD implementation. JAX-EPBD efficiently captures time-dependent behaviors, including bubble lifetimes and base flipping kinetics, enabling genome-scale analyses. Applying it to transcription factor (TF) binding affinity prediction using SELEX datasets, we observed consistent improvements in R2 values when incorporating breathing features with sequence data. Validating on the 77-bp AAV P5 promoter, JAX-EPBD revealed sequence-specific differences in bubble dynamics correlating with transcriptional activity. These findings establish JAX-EPBD as a powerful and scalable tool for understanding DNA breathing dynamics and their role in gene regulation and transcription factor binding.

bioinformatics↗

Scalable DNA Feature Generation and Transcription Factor Binding Prediction via Deep Surrogate Models

Simulating DNA breathing dynamics, for instance Extended Peyrard-Bishop-Dauxois (EPBD) model, across the entire human genome using traditional biophysical methods like pyDNA-EPBD is computationally prohibitive due to intensive techniques such as Markov Chain Monte Carlo (MCMC) and Langevin dynamics. To overcome this limitation, we propose a deep surrogate generative model utilizing a conditional Denoising Diffusion Probabilistic Model (DDPM) trained on DNA sequence-EPBD feature pairs. This surrogate model efficiently generates high-fidelity DNA breathing features conditioned on DNA sequences, reducing computational time from months to hours-a speedup of over 1000 times. By integrating these features into the EPBDxDNABERT-2 model, we enhance the accuracy of transcription factor (TF) binding site predictions. Experiments demonstrate that the surrogate-generated features perform comparably to those obtained from the original EPBD framework, validating the models efficacy and fidelity. This advancement enables real-time, genome-wide analyses, significantly accelerating genomic research and offering powerful tools for disease understanding and therapeutic development.

genomics↗

Melanopsin ganglion cells in the mouse retina independently evoke pupillary light reflex

PurposeThe pupillary light reflex (PLR) is crucial for protecting the retina from bright light. The intrinsic photosensitive ganglion cells (ipRGCs) in the retina mediate the PLR, which directly sense light and receive inputs from rod/cone photoreceptors. Previous work used genetic knockout mice to reveal that rod/cone photoreceptors drive transient constriction, and ipRGCs drive the sustained component. We acutely ablated photoreceptors by a chemical injection to examine the role of rod and cone photoreceptors in PLR. MethodsPLR and the multiple electrode array (MEA) recording were conducted with C57BL6/J (wildtype: WT) and Cnga3-/-; Gnat1-/- (rod/cone dysfunctional) mice. n-Nitroso-n-methylurea (MNU) was applied to C57 mice by intraperitoneal injection, and PLR was conducted after 5-7 days of injection. Three different light levels (mesopic, low photopic, and high photopic) were tested. Immunohistochemistry was conducted using the anti-Gnat1 and anti-melanopsin antibodies with DAPI. ResultsPLR was induced by all light levels we tested, and the level of constriction increased as the light level increased. After the MNU injection, PLR was not induced at mesopic light stimulus, but was fully induced by high light. The level of PLR was identical between WT and MNU mice, suggesting that ipRGCs fully contributed to the PLR at this light level. Immunohistochemistry revealed that photoreceptors were ablated by the MNU injection, but ipRGCs were preserved. The MEA recording revealed that a population of ipRGCs generated fast and robust spikes in MNU-injected retinal tissues in ex vivo. ConclusionsContrary to previous observations, our results demonstrate that ipRGCs are the major contributor to the PLR induced by high light.

neuroscience↗

Advancing Transcription Factor Binding Site Prediction Using DNA Breathing Dynamics and Sequence Transformers via Cross Attention

Understanding the impact of genomic variants on transcription factor binding and gene regulation remains a key area of research, with implications for unraveling the complex mechanisms underlying various functional effects. Our study delves into the role of DNAs biophysical properties, including thermodynamic stability, shape, and flexibility in transcription factor (TF) binding. We developed a multi-modal deep learning model integrating these properties with DNA sequence data. Trained on ChIP-Seq (chromatin immunoprecipitation sequencing) data in vivo involving 690 TF-DNA binding events in human genome, our model significantly improves prediction performance in over 660 binding events, with up to 9.6% increase in AUROC metric compared to the baseline model when using no DNA biophysical properties explicitly. Further, we expanded our analysis to in vitro high-throughput Systematic Evolution of Ligands by Exponential enrichment (SELEX) and Protein Binding Microarray (PBM) datasets, comparing our model with established frameworks. The inclusion of DNA breathing features consistently improved TF binding predictions across different cell lines in these datasets. Notably, for complex ChIP-Seq datasets, integrating DNABERT2 with a cross-attention mechanism provided greater predictive capabilities and insights into the mechanisms of disease-related non-coding variants found in genome-wide association studies. This work highlights the importance of DNA biophysical characteristics in TF binding and the effectiveness of multi-modal deep learning models in gene regulation studies.

bioinformatics↗

Examining DNA Breathing with pyDNA-EPBD

MotivationThe two strands of the DNA double helix locally and spontaneously separate and recombine in living cells due to the inherent thermal DNA motion.This dynamics results in transient openings in the double helix and is referred to as "DNA breathing" or "DNA bubbles." The propensity to form local transient openings is important in a wide range of biological processes, such as transcription, replication, and transcription factors binding. However, the modeling and computer simulation of these phenomena, have remained a challenge due to the complex interplay of numerous factors, such as, temperature, salt content, DNA sequence, hydrogen bonding, base stacking, and others. ResultsWe present pyDNA-EPBD, a parallel software implementation of the Extended Peyrard-Bishop-Dauxois (EPBD) nonlinear DNA model that allows us to describe some features of DNA dynamics in detail. The pyDNA-EPBD generates genomic scale profiles of average base-pair openings, base flipping probability,DNA bubble probability, and calculations of the characteristically dynamic length indicating the number of base pairs statistically significantly affected by a single point mutation using the Markov Chain Monte Carlo (MCMC) algorithm.

genomics↗

GOProFormer: A Multi-modal Transformer Method for Gene Ontology Protein Function Prediction

Protein Language Models (PLMs) are shown capable of learning sequence representations useful for various prediction tasks, from subcellular localization, evolutionary relationships, family membership, and more. They have yet to be demonstrated useful for protein function prediction. In particular, the problem of automatic annotation of proteins under the Gene Ontology (GO) framework remains open. This paper makes two key contributions. It debuts a novel method that leverages the transformer architecture in two ways. A sequence transformer encodes protein sequences in a task-agnostic feature space. A graph transformer learns a representation of GO terms while respecting their hierarchical relationships. The learned sequence and GO terms representations are combined and utilized for multi-label classification, with the labels corresponding to GO terms. The method is shown superior over recent representative GO prediction methods. The second major contribution in this paper is a deep investigation of different ways of constructing training and testing datasets. The paper shows that existing approaches under- or over-estimate the generalization power of a model. A novel approach is proposed to address these issues, resulting a new benchmark dataset to rigorously evaluate and compare methods and advance the state-of-the-art.

bioinformatics↗