bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.09.24.754079

Bridging the gap between omics and structural data: A framework for interpreting protein-RNA interaction specificity

Abstract

Protein-RNA interactions play a central role in many cellular processes, such as gene regulation, protein synthesis and viral infections. Although a few thousand protein-RNA complexes have been structurally characterized, they represent only a small fraction of all interactions. This lack of experimental data limits the ability of deep-learning based approaches such as AlphaFold3 to predict structures of protein-RNA interactions. In this study, we investigate protein-RNA binding specificity by enriching experimental structures with omics data. To this end, we present a scoring approach and associated web resource quantifying the agreement between experimental structures of protein-RNA interfaces and their omics-derived binding preferences, allowing us to identify interaction motif cores. Through key structural and evolutionary features, we further highlight that these motif cores correspond to important interface regions. We then leverage the dataset of protein-RNA complexes for which structural information can be combined with binding preferences in order to benchmark the ability of AlphaFold3 to predict protein-RNA interaction specificity. To this aim, we run AlphaFold3 predictions using different RNA inputs, from the exact sequence present in the experimental structure to a non-specific sequence, including sequences embedding the consensus binding motif from in vitro experiments. We show that despite good prediction quality, the sensitivity of AlphaFold3 to the exact RNA sequence used as input indicates signs of memorization. We examine some particular complexes and uncover challenges encountered by our scoring workflow and AlphaFold3, especially in the case of alternative binding modes. Altogether, this work evidences both the promise and current limitations of deep learning approaches for protein-RNA structure prediction, and provides a resource to guide their further development.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fauconnet, Y., Deng, S., Koundri, R., Pagani, M., Quignot, C., Antonio, D., Andreani, J.. 2026-09-30. Bridging the gap between omics and structural data: A framework for interpreting protein-RNA interaction specificity. https://doi.org/10.64898/2026.09.24.754079

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

A moving target: non-stationary selection governs unsupervised prediction of viral fitness

Anticipating how mutations change viral fitness is central to genomic surveillance and vaccine design, yet the supervised phenotype data behind the most accurate variant-effect predictors are unavailable for most emerging pathogens. We ask how far label-free scoring can go using only sequences, their evolutionary history, and structure. We assemble a modular, fully unsupervised pipeline that estimates a few interpretable terms (intrinsic replicative fitness, antigenic escape, and realized growth), and that lets each term be produced by more than one estimator, so the estimator itself becomes a testable modeling choice. Benchmarking the intrinsic term on 21 viral deep-mutational-scanning assays from ProteinGym, we find that a 650-million-parameter single-sequence protein language model predicts viral mutational fitness weakly and heterogeneously (mean Spearman 0.15), whereas a trivial site-independent alignment model more than doubles it (0.39, better on 17 of 21 assays), with the largest gains on the antigenic surface proteins where the language model fails. Yet the ordering reverses across 186 non-viral ProteinGym assays, where the language model instead exceeds the alignment model, localizing the weakness to viral families under-represented in the model's training data. Alignment-conditioned language models (MSA Transformer, Tranception) recover this accuracy but do not clearly exceed the simple alignment, so the decisive feature is the family alignment, not model scale or architecture. Our central result is evolutionary. Using dated samples of SARS-CoV-2 spike and influenza H3N2 hemagglutinin, we show that the epoch of the alignment is itself a leading, virus-specific determinant of accuracy. This traces to non-stationary selection: the site-specific amino-acid preferences drift over time, abruptly for spike at the emergence of Omicron and gradually for H3N2 hemagglutinin. A phylogenetic mutation-selection estimator does not match the far cheaper alignment model, falling significantly below it on matched data. Unsupervised viral fitness prediction is, then, as much an evolutionary problem as a modeling one.

bioinformatics↗

Chikungunya Virus Infection of Monocytes Generates a Persistent Macrophage Reservoir That Promotes IFN-π/IL-27-STAT1- and NF-κB-Driven Chronic Arthritis

Chikungunya virus (CHIKV) infection causes acute febrile illness that frequently progresses to chronic arthralgia. Although monocytes are CHIKV targets, whether viral exposure drives their differentiation into macrophages to sustain persistent inflammation remains incompletely understood. To address this, we integrated transcriptomic profiling of pediatric CHIKV patients during acute and convalescent phases, in vitro differentiation of primary human monocytes into macrophages under CHIKV infection, and single-cell RNA-sequencing (scRNA-seq) of joint-associated macrophages from an immunocompetent mouse model of chronic CHIKV arthritis at 28 days post-infection. In acute patients, we observed a marked upregulation of monocyte/macrophage markers alongside robust, IFN-dependent JAK-STAT signaling. In vitro, CHIKV-infected monocytes differentiated into persistently infected macrophages (MDM-CHIKV) that closely aligned with an M1-polarized transcriptional profile. These cells exhibited STAT1/NF-{kappa}B-dependent inflammation, ISG induction, and enhanced antigen presentation capacity. Remarkably, MDM-CHIKV uniquely produced IFN-{pi}/IL-27, revealing an alternative non-canonical axis where JAK-STAT activation persists despite the absence of classical type I IFNs. Functionally, MDM-CHIKV displayed an enhanced respiratory burst upon immune complex uptake, pointing to Fc{gamma}R-mediated pathology, as well as complement activation and T-cell stimulatory capacity. However, upon TLR4 challenge, these cells retained partial regulatory responses via IL-10 and TGF-{beta}; this hybrid immunophenotype offers a potential mechanistic explanation for relapsing arthralgia. Additionally, scRNA-seq of murine joints confirmed that persistent CHIKV replication was restricted to a distinct macrophage subset that resembles MDM-CHIKV, establishing these cells as the primary in vivo viral reservoir. Mirroring our human findings, joint-associated murine macrophage subsets exhibited elevated IFN-{pi}/IL-27 expression alongside robust NF-{kappa}B/STAT1-dependent inflammatory signatures (Tnf, Il1{beta}, Il6, Ccl2, Ccl5, Cxcl9, Cxcl10, Cxcl1-3, and/or Ptgs2), underscoring their critical role in sustaining chronic arthritis. Collectively, our findings demonstrate that CHIKV actively drives monocyte-to-macrophage differentiation into a persistently infected, M1-skewed hybrid phenotype that serves as a viral reservoir and fuels pro-arthritic inflammation. Consequently, macrophage reprogramming, the non-canonical IFN-{pi}/IL-27-STAT1 axis, Fc{gamma}R-mediated phagocytosis, and complement activation emerge as central pathogenic nodes and promising therapeutic targets for mitigating chronic post-CHIKV arthropathy.

bioinformatics↗

ProxiNet transfers spatially learned cellular proximity to dissociated single-cell transcriptomes

Spatial transcriptomics reveals cellular organization within intact tissues, whereas dissociated single-cell RNA sequencing provides broad transcriptomic coverage but loses information about cellular proximity and neighborhood structure. Here, we developed ProxiNet, a spatially supervised framework that learns transcriptomic signatures of pairwise cellular proximity from spatial reference datasets and transfers these relationships to dissociated single-cell transcriptomes. ProxiNet predicted cellular proximity across brain regions and spatial technologies, including zero-shot cross-technology transfer, and gradient-based attribution identified genes and broader transcriptional programs associated with proximity predictions. In spatial datasets with known coordinates, ProxiNet-derived cellular neighborhoods recovered reproducible tissue organization and anatomical structure, providing independent spatial validation of the inferred proximity relationships. Applying the spatially calibrated model to dissociated scRNA-seq revealed heterogeneous cellular neighborhoods with distinct cell-type compositions and candidate communication programs. In an Alzheimer dataset, ProxiNet further identified age-associated remodeling of inferred neighborhoods, including an AD-associated neighborhood at 8 months characterized by distinct astrocyte and neuronal transcriptional states and candidate intercellular communication programs. Together, these results establish pairwise cellular proximity as an interpretable and transferable representation for extending spatially learned tissue organization to dissociated single-cell transcriptomes.

bioinformatics↗