bioRxiv ScienceSearch

Biology subjects

Madani Tonekaboni, S. A.

Publications and source records attributed to Madani Tonekaboni, S. A..

2 recordsLinked to original sources

CREAM: Clustering of genomic REgions Analysis Method

Cellular identity relies on cell type-specific gene expression profiles controlled by cis-regulatory elements (CREs), such as promoters, enhancers and anchors of chromatin interactions. CREs are unevenly distributed across the genome, giving rise to distinct subsets such as individual CREs and Clusters Of cis-Regulatory Elements (COREs), also known as super-enhancers. Identifying COREs is a challenge due to technical and biological features that entail variability in the distribution of distances between CREs within a given dataset. To address this issue, we developed a new unsupervised machine learning approach termed Clustering of genomic REgions Analysis Method (CREAM) that outperforms the Ranking Of Super Enhancer (ROSE) approach. Specifically CREAM identified COREs are enriched in CREs strongly bound by master transcription factors according to ChIP-seq signal intensity, are proximal to highly expressed genes, are preferentially found near genes essential for cell growth and are more predictive of cell identity. Moreover, we show that CREAM enables subtyping primary prostate tumor samples according to their CORE distribution across the genome. We further show that COREs are enriched compared to individual CREs at TAD boundaries and these are preferentially bound by CTCF and factors of the cohesin complex (e.g.: RAD21 and SMC3). Finally, using CREAM against transcription factor ChIP-seq reveals CTCF and cohesin-specific COREs preferentially at TAD boundaries compared to intra-TADs. CREAM is available as an open source R package (https://CRAN.R-project.org/package=CREAM) to identify COREs from cis-regulatory annotation datasets from any biological samples.

bioinformatics

Similarity identification in gene expression patterns as a new approach in phenotype classification

Stratifying healthy and malignant phenotypes and identifying their biological states using high-throughput molecular data has been the focus of many computational approaches during the last decade. Using multivariate changes in expression of genes within biological pathways, as fingerprints of complex phenotypes, we developed a new methodology for Similarity Identification in Gene expressioN (SIGN). In this approach, we use centroid classifier to identify phenotype of each biological sample. To obtain similarity of a given biological sample with classes of phenotypes, we defined a new distance measure, transcriptional similarity coefficient (TSC) which captures similarity of gene expression patterns between a biological pathway in two samples or populations. We showed that TSC, as an interpretable and stable distance measure in SIGN, captures all oncogenic hallmarks for breast cancer even with low sample size, by comparing healthy and patient tumor samples in the largest breast cancer dataset. In this study, we demonstrate that SIGN is a flexible, yet robust approach for classification based on transcriptomics data. Comparing early and late relapses within each molecular subtypes of breast cancer, our method enabled subtype-specific stratification of breast cancer patients into groups with significantly different survival. Moreover, we used SIGN to classify with more than 99% specificity the site of extraction of healthy and tumor samples from the Genotype-Tissue Expression (GTEx) and The Cancer Genome Atlas (TCGA) datasets. We showed that SIGN also enables robust identification of hematopoietic stem cell and progenitors within the hematopoietic hierarchy. We further explored chemical perturbation data in the Connectivity Map (CMAP) database and showed that SIGN was able to classify seven classes of drugs based on their mechanism of action. In conclusion, we showed that SIGN can be used to achieve interpretable and robust transcriptomic-based classification of healthy and malignant samples, as well as drugs based on their known mechanism of action, supporting the generalizability and relevance of the method for the analysis of gene expression profiles.

bioinformatics