bioRxiv ScienceSearch

Biology subjects

Modolo, L.

Publications and source records attributed to Modolo, L..

3 recordsLinked to original sources

Interplay between coding and exonic splicing regulatory sequences

The inclusion of exons during the splicing process depends on the binding of splicing factors to short low-complexity regulatory sequences. The relationship between exonic splicing regulatory sequences and coding sequences is still poorly understood. We demonstrate that exons that are coregulated by any splicing factor share a similar nucleotide composition bias. We next demonstrate that coregulated exons preferentially code for amino acids with similar physicochemical properties because of the non-randomness properties of the genetic code. Indeed, amino acids sharing physicochemical properties correspond to codons that have the same nucleotide composition bias. These observations reveal an unanticipated bidirectional interplay between the physicochemical features encoded by exons and exon splicing regulation by splicing factors. We propose that the splicing regulation of an exon by a splicing factor is tightly interconnected with the physicochemical properties of the exon-encoded protein domain depending on the splicing-factor affinity for specific nucleotides.

genomics

Condensin controls cellular RNA levels through the accurate segregation of chromosomes instead of directly regulating transcription

Condensins are genome organisers that shape chromosomes and promote their accurate transmission. Several studies have also implicated condensins in gene expression, although the mechanisms have remained enigmatic. Here, we report on the role of condensin in gene expression in fission and budding yeasts. In contrast to previous studies, we provide compelling evidence that condensin plays no direct role in the maintenance of the transcriptome, neither during interphase nor during mitosis. We further show that the changes in gene expression in post-mitotic fission yeast cells that result from condensin inactivation are largely a consequence of chromosome missegregation during anaphase, which notably depletes the RNA-exosome from daughter cells. Crucially, preventing karyotype abnormalities in daughter cells restores a normal transcriptome despite condensin inactivation. Thus, chromosome instability, rather than a direct role of condensin in the transcription process, changes gene expression. This knowledge challenges the concept of gene regulation by canonical condensin complexes.

molecular biology

Probabilistic Count Matrix Factorization for Single Cell Expression Data Analysis

The development of high throughput single-cell technologies now allows the investigation of the genome-wide diversity of transcription. This diversity has shown two faces: the expression dynamics (gene to gene variability) can be quantified more accurately, thanks to the measurement of lowly-expressed genes. Second, the cell-to-cell variability is high, with a low proportion of cells expressing the same gene at the same time/level. Those emerging patterns appear to be very challenging from the statistical point of view, especially to represent and to provide a summarized view of single-cell expression data. PCA is one of the most powerful frameworks to provide a suitable representation of high dimensional datasets, by searching for new axes catching the most variability in the data. Unfortunately, classical PCA is based on Euclidean distances and projections that work poorly in presence of over-dispersed counts showing zero-inflation. We propose a probabilistic Count Matrix Factorization (pCMF) approach for single-cell expression data analysis, that relies on a sparse Gamma-Poisson factor model. This hierarchical model is inferred using a variational EM algorithm. We show how this probabilistic framework induces a geometry that is suitable for single-cell data, and produces a compression of the data that is very powerful for clustering purposes. Our method is competed to other standard representation methods like ZIFA and t-SNE, and we illustrate its performance on simulated and publicly available data.

bioinformatics