bioRxiv Science⌕ Search

Biology subjects

Prodic, S.

Publications and source records attributed to Prodic, S..

4 recordsLinked to original sources

ModkitOpt: Systematic optimisation of modkit parameters for accurate nanopore-based RNA modification detection

Despite enabling single-molecule detection of RNA modifications, nanopore direct RNA sequencing lacks standardised approaches for modification site calling. This requires accurately quantifying per-site modification stoichiometry and selecting an appropriate stoichiometry cutoff to classify sites. Here, we show that modkit, the de facto standard tool for estimating modification stoichiometry, is highly sensitive to parameter selection, and its heuristic parameter choice consistently leads to markedly sub-optimal site calling, which is exacerbated in datasets where dorado prediction confidence is heterogeneous. We also demonstrate that the choice of stoichiometry cutoff significantly affects false positive and false negative rates for called sites, leading to divergent biological conclusions. To address both limitations we introduce ModkitOpt, a pipeline that identifies then applies the optimal modkit parameters and stoichiometry cutoff for any modification type given a set of validated sites, producing optimised site calls for any nanopore sequencing dataset. Across multiple modification types and biological contexts, ModkitOpt consistently recovers the precision and recall of called sites, establishing a robust framework for standardised RNA modification stoichiometry estimation and site calling from nanopore direct RNA sequencing. ModkitOpt is available at https://github.com/comprna/modkitopt.

bioinformatics↗

SWARM: A Single-Molecule Workflow for High-Precision Profiling of RNA Modifications

Nanopore direct RNA sequencing promises to decode the epitranscriptome by detecting multiple modifications on individual RNA molecules, but its potential for biological discovery is hampered by high false-positive rates. We present SWARM, an AI-based framework designed to overcome this fundamental limitation. Its key innovation is a crosstalk-aware training strategy that incorporates non-target modifications and orthogonally validated cellular signals, enabling high-precision detection of m6A, pseudouridine ({Psi}), and m5C at single-nucleotide and single-molecule resolution. Using rigorous in vitro and cellular RNA benchmarks, SWARM outperforms existing tools and maintains strong agreement with orthogonal methods. Applying SWARM across mammalian tissues reveals thousands of novel modification sites with confirmed motifs and localisation patterns. Our high-resolution multi-tissue modification map revealed no evidence of widespread m6A-{Psi} interplay in predominant writer contexts, challenging models of a coordinated epitranscriptomic code. We further discovered a previously unrecognised splicing-shaped mode of {Psi} deposition, whereby TRUB1-mediated pseudouridylation preferentially occurs after exon-exon ligation, consistent with local RNA structure stabilisation. SWARM provides a robust, universally applicable tool for epitranscriptome discovery.

bioinformatics↗

Nano-Mod-Amp reveals RNA sequence, structural and cell type specific features of pseudouridylation by PUS7

Pseudouridines are abundant mRNA modifications that can impact splicing, translation, and stability to tune gene expression. PUS7 is one of the major mRNA pseudouridine synthase whose dysregulation leads to neurodevelopmental disorders and cancer, underscoring the critical function of PUS7-dependent pseudouridines. Beyond a short and degenerate consensus sequence, the molecular mechanisms underlying PUS7-mediated pseudouridylation remain unknown. A lack of targeted, high-throughput pseudouridine detection methods limits simultaneous interrogation of PUS7 regulatory features across many experimental conditions. We developed novel Nanopore sequencing tools, including Nano-Mod-Amp, to reveal pseudouridine stoichiometry, its RNA structural context, and dependence on PUS7 levels at specific sites across biological conditions. We identified a novel RNA structural signature that is associated with more efficient mRNA modification by PUS7. Pseudouridines are largely responsive to modulations in PUS7 protein levels, demonstrating the regulatory potential of varying PUS7 levels across cellular conditions. Conversely, PUS7 activity is also regulated in a cell-type specific manner, independent of PUS7 expression levels in a manner consistent with regulation by RNA structure and RNA binding proteins. Together, we developed Nanopore sequencing tools and uncovered new mechanisms of PUS7 regulation with a framework that can be applied to other RNA-modifying enzymes to query the regulation of the epitranscriptome. HighlightsO_LINanopore direct RNA sequencing identifies PUS7-dependent pseudouridines with stoichiometry. C_LIO_LINano-Mod-Amp quantifies PUS7-dependent pseudouridines at hundreds of sites in high-throughput. C_LIO_LIMPRAs define RNA sequence and structural features associated with modification by PUS7. C_LIO_LIIndividual PUS7 target pseudouridines are substoichiometric and poised for regulation. C_LIO_LIPUS7 activity is regulated by cell type in the absence of differences in PUS7 protein levels. C_LI

molecular biology↗

A basic framework governing splice-site choice in eukaryotes

Changes in splicing are observed between cells, tissues, organs, individuals, and species. These changes can mediate phenotypic variation ranging from flowering time differences in plants to genetic diseases in humans. However, the genomic determinants of splicing variation are largely unknown. Here, we quantified the usage of individual splice-sites and uncover extensive variation between individuals (genotypes) in Arabidopsis, Drosophila and Humans. We used this robust quantitative measure as a phenotype and mapped variation in splice-site usage using Genome-Wide Association Studies (GWAS). By carrying out more than 130,000 GWAS with splice-site usage phenotypes, we reveal genetic variants associated with differential usage of specific splice-sites. Our analysis conclusively shows that most of the common, genetically controlled variation in splicing is cis and there are no major trans hotspots in any of the three analyzed species. High-resolution mapping allowed us to determine genome-wide patterns that govern splice-site choice. We reveal that the variability in the intronic hexamer sequence (GT[N]4 or [N]4AG) differentiates intrinsic splice-site strength and is among the primary determinants of splice-site choice. Experimental analysis validates the primary role for intronic hexamer sequences in conferring splice-site decisions. Transcriptome analyses in diverse species across the tree of life reveals that hexamer rankings explains splice-site choices from yeast to plants to humans, forming the basic framework of the splicing code in eukaryotes.

genetics↗