bioRxiv ScienceSearch

Biology subjects

Rogan, P. K.

Publications and source records attributed to Rogan, P. K..

3 recordsLinked to original sources

Clustered, information-dense transcription factor binding sites identify genes with similar tissue-wide expression profiles

BackgroundThe distribution and composition of cis-regulatory modules (e.g. transcription factor binding site (TFBS) clusters) in promoters substantially determine gene expression patterns and TF targets, whose expression levels are significantly regulated by TF binding. TF knockdown experiments have revealed correlations between TF binding profiles and gene expression levels. We present a general framework capable of predicting genes with similar tissue-wide expression patterns from activated or repressed TF targets using machine learning to combine TF binding and epigenetic features.\n\nMethodsGenes with correlated expression patterns across 53 tissues were identified according to their Bray-Curtis similarity. DNase I HyperSensitive region (DHS) -accessible promoter intervals of direct TF target genes were scanned with previously derived information theory-based position weight matrices (iPWMs) of 82 TFs. Features from information density-based TFBS clusters were used to predict target genes with machine learning classifiers. The accuracy, specificity and sensitivity of the classifiers were determined for different feature sets. Mutations in TFBSs were also introduced to examine their impact on cluster densities and the regulatory states of predicted target genes.\n\nResultsWe initially chose the glucocorticoid receptor gene (NR3C1), whose regulation has been extensively studied, to test this approach. SLC25A32 and TANK were found to exhibit the most similar expression patterns to this gene across 53 tissues. Prediction of other genes with similar expression profiles was significantly improved by eliminating inaccessible promoter intervals based on DHSs. A Random Forest classifier exhibited the best performance in detecting such coordinately regulated genes (accuracy was 0.972 for training, 0.976 for testing). Target gene prediction was confirmed using CRISPR knockdown data of TFs, which was more accurate than siRNA inactivation. Mutation analyses of TFBSs also revealed that one or more information-dense TFBS clusters in promoters are required for accurate target gene prediction.\n\nConclusionsMachine learning based on TFBS information density, organization, and chromatin accessibility accurately identifies gene targets with comparable tissue-wide expression patterns. Multiple, information-dense TFBS clusters in promoters appear to protect promoters from the effects of deleterious binding site mutations in a single TFBS that would effectively alter the expression state of these genes.

bioinformatics

Predicting Response to Platin Chemotherapy Agents with Biochemically-inspired Machine Learning

Selection of effective genes that accurately predict chemotherapy response could improve cancer outcomes. We compare optimized gene signatures for cisplatin, carboplatin, and oxaliplatin response in the same cell lines, and respectively validate each with cancer patient data. Supervised support vector machine learning was used to derive gene sets whose expression was related to cell line GI50 values by backwards feature selection with cross-validation. Specific genes and functional pathways distinguishing sensitive from resistant cell lines are identified by contrasting signatures obtained at extreme vs. median GI50 thresholds. Ensembles of gene signatures at different thresholds are combined to reduce dependence on specific GI50 values for predicting drug response. The most accurate models for each platin are: cisplatin: BARD1, BCL2, BCL2L1, CDKN2C, FAAP24, FEN1, MAP3K1, MAPK13, MAPK3, NFKB1, NFKB2, SLC22A5, SLC31A2, TLR4, TWIST1; carboplatin: AKT1, EIF3K, ERCC1, GNGT1, GSR, MTHFR, NEDD4L, NLRP1, NRAS, RAF1, SGK1, TIGD1, TP53, VEGFB, VEGFC; oxaliplatin: BRAF, FCGR2A, IGF1, MSH2, NAGK, NFE2L2, NQO1, PANK3, SLC47A1, SLCO1B1, UGT1A1. TCGA bladder, ovarian and colorectal cancer patients were used to test cisplatin, carboplatin and oxaliplatin signatures (respectively), resulting in 71.0%, 60.2% and 54.5% accuracy in predicting disease recurrence and 59%, 61% and 72% accuracy in predicting remission. One cisplatin signature predicted 100% of recurrence in non-smoking bladder cancer patients (57% disease-free; N=19), and 79% recurrence in smokers (62% disease-free; N=35). This approach should be adaptable to other studies of chemotherapy response, independent of drug or cancer types.

pharmacology and toxicology

Accurate Cytogenetic Biodosimetry Through Automation Of Dicentric Chromosome Curation And Metaphase Cell Selection

Software to automate digital pathology relies on image quality and the rates of false positive and negative objects in these images. Cytogenetic biodosimetry detects dicentric chromosomes (DCs) that arise from exposure to ionizing radiation, and determines radiation dose received from the frequency of DCs. We present image segmentation methods to rank high quality cytogenetic images and eliminate suboptimal metaphase cell data based on novel quality measures. Improvements in DC recognition increase the accuracy of dose estimates, by reducing false positive (FP) DC detection. A set of chromosome morphology segmentation methods selectively filtered out false DCs, arising primarily from extended prometaphase chromosomes, sister chromatid separation and chromosome fragmentation. This reduced FPs by 55% and was highly specific to the abnormal structures ([&ge;]97.7%). Additional procedures were then developed to fully automate image review, resulting in 6 image-level filters that, when combined, selectively remove images with consistently unparsable or incorrectly segmented chromosome morphologies. Overall, these filters can eliminate half of the FPs detected by manual image review. Optimal image selection and FP DCs are minimized by combining multiple feature based segmentation filters and a novel image sorting procedure based on the known distribution of chromosome lengths. Applying the same segmentation filtering procedures to both calibration and test sample image data reduced the average dose estimation error from 0.4Gy to <0.2Gy, obviating the need to first manually review these images. This reliable and scalable solution enables batch processing multiple samples of unknown dose, and meets current requirements for triage radiation biodosimetry of high quality metaphase cell preparations.

genetics