bioRxiv ScienceSearch

Biology subjects

Andrade-Navarro, M. A.

Publications and source records attributed to Andrade-Navarro, M. A..

5 recordsLinked to original sources

Computational Chromosome Conformation Capture by Correlation of ChIP-seq at CTCF motifs

BackgroundKnowledge of the three-dimensional structure of the genome is necessary to understand how gene expression is regulated. Recent experimental techniques such as Hi-C or ChIA-PET measure long-range interactions genome-wide but are experimentally elaborate and have limited resolution. Here, we present Computational Chromosome Conformation Capture by Correlation of ChIP-seq at CTCF motifs (7C).\n\nResultsWhile ChIP-seq was not designed to detect contacts, the formaldehyde treatment in the ChIP-seq protocol cross-links proteins with each other and with DNA. Consequently, also regions that are not directly bound by the targeted TF but interact with the binding site via chromatin looping are co-immunoprecipitated and sequenced. This produces minor ChIP-seq signals at loop anchor regions close to the directly bound site. We use the position and shape of ChIP-seq signals around CTCF motif pairs to predict whether they interact or not.\n\nWe applied 7C to all CTCF motif pairs within 1 MB in the human genome and validated predicted interactions with high-resolution Hi-C and ChIA-PET. A single ChIP-seq experiment from known architectural proteins (CTCF, Rad21, Znf143) but also from other TFs (like TRIM22 or RUNX3) predicts loops accurately. Importantly, 7C predicts loops in cell types and for TF ChIP-seq datasets not used in training.\n\nConclusion7C predicts chromatin loops with base-pair resolution and can be used to associate TF binding sites to regulated genes in a condition-specific manner. Furthermore, profiling of hundreds of ChIP-seq datasets results in novel candidate factors functionally involved in chromatin looping. Our method is available as an R package: https://ibn-salem.github.io/sevenC/

genomics

Evaluating Cell Identity from Transcription Profiles

Induced pluripotent stem cells (iPS) and direct lineage programming offer promising autologous and patient-specific sources of cells for personalized drug-testing and cell-based therapy. Before these engineered cells can be widely used, it is important to evaluate how well the engineered cell types resemble their intended target cell types. We have developed a method to generate CellScore, a cell identity score that can be used to evaluate the success of an engineered cell type in relation to both its initial and desired target cell type, which are used as references. Of 20 cell transitions tested, the most successful transitions were the iPS cells (CellScore > 0.9), while other transitions (e.g. induced hepatocytes or motor neurons) indicated incomplete transitions (CellScore < 0.5). In principle, the method can be applied to any engineered cell undergoing a cell transition, where transcription profiles are available for the reference cell types and the engineered cell type.\n\nHighlightsO_LIA curated standard dataset of transcription profiles from normal cell types was created.\nC_LIO_LICellScore evaluates the cell identity of engineered cell types, using the curated dataset.\nC_LIO_LICellScore considers the initial and desired target cell type.\nC_LIO_LICellScore identifies the most successfully engineered clones for further functional testing.\nC_LI

bioinformatics

Evolutionary stability of topologically associating domains is associated with conserved gene regulation

BackgroundThe human genome is highly organized in the three-dimensional nucleus. Chromosomes fold locally into topologically associating domains (TADs) defined by increased intra-domain chromatin contacts. TADs contribute to gene regulation by restricting chromatin interactions of regulatory sequences, such as enhancers, with their target genes. Disruption of TADs can result in altered gene expression and is associated to genetic diseases and cancers. However, it is not clear to which extent TAD regions are conserved in evolution and whether disruption of TADs by evolutionary rearrangements can alter gene expression.\n\nResultsHere, we hypothesize that TADs represent essential functional units of genomes, which are selected against rearrangements during evolution. We investigate this using whole-genome alignments to identify evolutionary rearrangement breakpoints of different vertebrate species. Rearrangement breakpoints are strongly enriched at TAD boundaries and depleted within TADs across species. Furthermore, using gene expression data across many tissues in mouse and human, we show that genes within TADs have more conserved expression patterns. Disruption of TADs by evolutionary rearrangements is associated with changes in gene expression profiles, consistent with a functional role of TADs in gene expression regulation.\n\nConclusionsTogether, these results indicate that TADs are conserved building blocks of genomes with regulatory functions that are often reshuffled as a whole instead of being disrupted by rearrangements.

genomics

The latent geometry of the human protein interaction network

To mine valuable information from the complex architecture of the human protein interaction network (hPIN), we require models able to describe its growth and dynamics accurately. Here, we present evidence that uncovering the latent geometry of the hPIN can ease challenging problems in systems biology. We embedded the hPIN to hyperbolic space, whose geometric properties reflect the characteristic scale invariance and strong clustering of the network. Interestingly, the inferred hyperbolic coordinates of nodes capture biologically relevant features, like protein age, function and cellular localisation. We also realised that the shorter the distance between two proteins in the embedding space, the higher their connection probability, which resulted in the prediction of plausible protein interactions. Finally, we observed that proteins can efficiently communicate with each other via a greedy routeing process, guided by the latent geometry of the hPIN. When analysed from the appropriate biological context, these efficient communication channels can be used to determine the core members of signal transduction pathways and to study how system perturbations impact their efficiency.

systems biology

A reliable and unbiased human protein network with the disparity filter

The living cell operates thanks to an intricate network of protein interactions. Proteins activate, transport, degrade, stabilise and participate in the production of other proteins. As a result, a reliable and systematically generated protein wiring diagram is crucial for a deeper understanding of cellular functions. Unfortunately, current human protein networks are noisy and incomplete. Also, they suffer from both study and technical biases: heavily studied proteins (e.g. those of pharmaceutical interest) are known to be involved in more interactions than proteins described in only a few publications. Here, we use the experimental evidence supporting the interaction between proteins, in conjunction with the so-called disparity filter, to construct a reliable and unbiased proteome-scale human interactome. The application of a global filter, i.e. only considering interactions with multiple pieces of evidence, would result in an excessively pruned network. In contrast, the disparity filter preserves interactions supported by a statistically significant number of studies and does not overlook small-scale protein associations. The resulting disparity-filtered protein network covers 67% of the human proteome and retains most of the networks weight and connectivity properties.

bioinformatics