bioRxiv Science⌕ Search

Biology subjects

Baulin, E. F.

Publications and source records attributed to Baulin, E. F..

3 recordsLinked to original sources

Statistical framework for calling allelic imbalance in high-throughput sequencing data

High-throughput sequencing facilitates large-scale studies of gene regulation and allows tracing the associations of individual genomic variants with changes in gene expression. Compared to classic association studies, allelic imbalance at heterozygous variants captures the functional effects of the regulatory genome variation with smaller sample sizes and higher sensitivity. Yet, the identification of allele-specific events from allelic read counts remains non-trivial due to multiple sources of technical and biological variability, which induce data-dependent biases and overdispersion. Here we present MIXALIME, a novel computational framework for calling allele-specific events in diverse omics data with a repertoire of statistical models accounting for read mapping bias and copy-number variation. We benchmark MIXALIME against existing tools and demonstrate its practical usage by constructing an atlas of allele-specific chromatin accessibility, UDACHA, from thousands of available datasets obtained from diverse cell types. Availabilityhttps://github.com/autosome-ru/MixALime, https://udacha.autosome.org

bioinformatics↗

SQUARNA - an RNA secondary structure prediction method based on a greedy stem formation model

Non-coding RNAs play a diverse range of roles in various cellular processes, with their spatial structure being pivotal to their function. The RNAs secondary structure is a key determinant of its overall fold. Given the scarcity of experimentally determined RNA 3D structures, understanding the secondary structure is vital for discerning the molecules function. Currently, there is no universally effective solution for de novo RNA secondary structure prediction. Existing methods are becoming increasingly complex without marked improvements in accuracy, and they often overlook critical elements such as pseudoknots. In this work, we introduce SQUARNA, a novel approach to de novo RNA secondary structure prediction. This method utilizes a simple, greedy stem formation model, addressing many of the limitations inherent in previous tools. Our benchmarks demonstrate that SQUARNA matches the performance of leading methods for single sequence inputs and significantly surpasses existing tools when applied to sequence alignment inputs.

bioinformatics↗

A comprehensive survey of long-range tertiary interactions and motifs in non-coding RNA structures

Understanding the 3D structure of RNA is key to understanding RNA function. RNA 3D structure is modular and can be seen as a composition of building blocks of various sizes called tertiary motifs. Currently, long-range motifs formed between distant loops and helical regions are largely less studied than the local motifs determined by the RNA secondary structure. We surveyed long-range tertiary interactions and motifs in a non-redundant set of non-coding RNA 3D structures. A new dataset of annotated LOng-RAnge RNA 3D modules (LORA) was built using an approach that does not rely on the automatic annotations of non-canonical interactions. An original algorithm, ARTEM, was developed for annotation-, sequence- and topology-independent superposition of two arbitrary RNA 3D modules. The proposed methods allowed us to identify and describe the most common long-range RNA tertiary motifs. Three basic interaction types were identified to be recurrent in the long-range RNA 3D modules: ribose-ribose interactions, canonical Type I and Type II A-minor interactions, and previously undescribed staple interactions. These three interaction types were found to be different building blocks of the same complex staple motifs common to non-coding RNA 3D structures.

bioinformatics↗