bioRxiv Science⌕ Search

Biology subjects

Joh, R. I.

Publications and source records attributed to Joh, R. I..

2 recordsLinked to original sources

MAGE: Monte Carlo method for Aberrant Gene Expression

Identifying genes that are aberrantly expressed is an important first step in the diagnosis and treatment of many diseases. Conventionally, differential expression (DE) analysis is used to screen gene expression profiles to identify functionally associated genes. DE often relies on the variance and fold change in expression from individual genes, which does not consider the expression of all other genes within the profile. When the overall gene expression is skewed, DE does not capture outliers in gene expression. To address this, we have developed a non-parametric DE method based on the probability density for an entire expression profile to select genes that deviate from the global distribution between two gene expression profiles with multiple replicates. Rather than assuming a particular distribution of expression per gene, our method assumes that aberrantly expressed genes (AEGs) will exhibit expression patterns distinguishable from non-AGEs which make up the majority of the profile. Here we introduce our nonparametric method (MAGE: Monte Carlo method for aberrant gene expression) and demonstrate that MAGE can identify AEGs that are not found by conventional DE analyses. The main feature of MAGE is (1) identifying outliers based on the expression profile of all genes rather than performing DE analyses on a per-gene basis and (2) consideration of the variance in expression between two different conditions. We also compared our results with traditional DE analysis as well as density-based clustering methods. MAGE produces consistent results in a variety of conditions and performs conservatively with the addition of noise. We also applied MAGE to single-cell RNA-seq samples and demonstrated that the analysis is robust with subsampling.

bioinformatics↗

Role of multiple pericentromeric repeats on heterochromatin assembly

Although the length and constituting sequences for pericentromeric repeats are highly variable across eukaryotes, the presence of multiple pericentromeric repeats is one of the conserved features of the eukaryotic chromosomes. Pericentromeric heterochromatin is often misregulated in human diseases, with the expansion of pericentromeric repeats in human solid cancers. In this article, we have developed a mathematical model of the RNAi-dependent methylation of H3K9 in the pericentromeric region of fission yeast. Our model, which takes copy number as an explicit parameter, predicts that the pericentromere is silenced only if there are many copies of repeats. It becomes bistable or desilenced if the copy number of repeats is reduced. This suggests that the copy number of pericentromeric repeats alone can determine the fate of heterochromatin silencing in fission yeast. Through sensitivity analysis, we identified parameters that favor bistability and desilencing. Stochastic simulation shows that faster cell division and noise favor the desilenced state. These results show the unexpected role of pericentromeric repeat copy number in gene silencing and elucidate how the copy number of silenced genomic regions may impact genome stability. Author SummaryPericentromeric repeats vary in length and sequences, but their presence is a conserved feature of eukaryotes. This suggests that the repetitive nature of pericentromeric sequences is an evolutionarily conserved feature of centromeres, which is under selective pressure. Here we developed a quantitative model for gene silencing at the fission yeast pericentromeric repeats. Our model is one of the first models which incorporates the copy number of pericentromeric repeats and predicts that the number of repeats can solely govern the dynamics of pericentromeric gene silencing. Our results suggest that the repeat copy number is a dynamic parameter for gene silencing, and copy-number-dependent silencing is an effective machinery to repress the repetitive part of the genome.

systems biology↗