bioRxiv ScienceSearch

Biology subjects

Vera, D.

Publications and source records attributed to Vera, D..

5 recordsLinked to original sources

Identification of cis elements for spatio-temporal control of DNA replication

The temporal order of DNA replication (replication timing, RT) is highly coupled with genome architecture, but cis-elements regulating spatio-temporal control of replication have remained elusive. We performed an extensive series of CRISPR mediated deletions and inversions and high-resolution capture Hi-C of a pluripotency associated domain (DppA2/4) in mouse embryonic stem cells. Whereas CTCF mediated loops and chromatin domain boundaries were dispensable, deletion of three intra-domain prominent CTCF-independent 3D contact sites caused a domain-wide delay in RT, shift in sub-nuclear chromatin compartment and loss of transcriptional activity, These \"early replication control elements\" (ERCEs) display prominent chromatin features resembling enhancers/promoters and individual and pair-wise deletions of the ERCEs confirmed their partial redundancy and interdependency in controlling domain-wide RT and transcription. Our results demonstrate that discrete cis-regulatory elements mediate domain-wide RT, chromatin compartmentalization, and transcription, representing a major advance in dissecting the relationship between genome structure and function.\n\nHighlightsO_LIcis-elements (ERCEs) regulate large scale chromosome structure and function\nC_LIO_LIMultiple ERCEs cooperatively control domain-wide replication\nC_LIO_LIERCEs harbor prominent active chromatin features and form CTCF-independent loops\nC_LIO_LIERCEs enable genetic dissection of large-scale chromosome structure-function.\nC_LI

molecular biology

Allele-specific control of replication timing and genome organization during development

DNA replication occurs in a defined temporal order known as the replication-timing (RT) program. RT is regulated during development in discrete chromosomal units, coordinated with transcriptional activity and 3D genome organization. Here, we derived distinct cell types from F1 hybrid musculus X castaneus mouse crosses and exploited the high single nucleotide polymorphism (SNP) density to characterize allelic differences in RT (Repli-seq), genome organization (Hi-C and promoter-capture Hi-C), gene expression (nuclear RNA-seq) and chromatin accessibility (ATAC-seq). We also present HARP: a new computational tool for sorting SNPs in phased genomes to efficiently measure allele-specific genome-wide data. Analysis of 6 different hybrid mESC clones with different genomes (C57BL/6, 129/sv and CAST/Ei), parental configurations and gender revealed significant RT asynchrony between alleles across ~12 % of the autosomal genome linked to sub-species genomes but not to parental origin, growth conditions or gender. RT asynchrony in mESCs strongly correlated with changes in Hi-C compartments between alleles but not SNP density, gene expression, imprinting or chromatin accessibility. We then tracked mESC RT asynchronous regions during development by analyzing differentiated cell types including extraembryonic endoderm stem (XEN) cells, 4 male and female primary mouse embryonic fibroblasts (MEFs) and neural precursors (NPCs) differentiated in vitro from mESCs with opposite parental configurations. Surprisingly, we found that RT asynchrony and allelic discordance in Hi-C compartments seen in mESCs was largely lost in all differentiated cell types, coordinated with a more uniform Hi-C compartment arrangement, suggesting that genome organization of homologues converges to similar folding patterns during cell fate commitment.

cell biology

iSeg: an efficient algorithm for segmentation of genomic and epigenomic data

BackgroundIdentification of functional elements of a genome often requires dividing a sequence of measurements along a genome into segments where adjacent segments have different properties, such as different mean values. This problem is often called the segmentation problem in the field of genomics, and the change-point problem in other scientific disciplines. Despite dozens of algorithms developed to address this problem in genomics research, methods with improved accuracy and speed are still needed to effectively tackle both existing and emerging genomic and epigenomic segmentation problems.\n\nResultsWe designed an efficient algorithm, called iSeg, for segmentation of genomic and epigenomic profiles. iSeg first utilizes dynamic programming to identify candidate segments and test for significance. It then uses a novel data structure based on two coupled balanced binary trees to detect overlapping significant segments and update them simultaneously during searching and refinement stages. Refinement and merging of significant segments are performed at the end to generate the final set of segments. By using an objective function based on the p-values of the segments, the algorithm can serve as a general computational framework to be combined with different assumptions on the distributions of the data. As a general segmentation method, it can segment different types of genomic and epigenomic data, such as DNA copy number variation, nucleosome occupancy, nuclease sensitivity, and differential nuclease sensitivity data. Using simple t-tests to compute p-values across multiple datasets of different types, we evaluate iSeg using both simulated and experimental datasets and show that it performs satisfactorily when compared with some other popular methods, which often employ more sophisticated statistical models. Implemented in C++, iSeg is also very computationally efficient, well suited for large numbers of input profiles and data with very long sequences.\n\nConclusionsWe have developed an effective and efficient general-purpose segmentation tool for sequential data and illustrated its use in segmentation of genomic and epigenomic profiles.

bioinformatics

SRSF shape analysis for sequencing data reveal newdifferentiating patterns

MotivationSequencing-based methods to examine fundamental features of the genome, such as gene expression and chromatin structure, rely on inferences from the abundance and distribution of reads derived from Illumina sequencing. Drawing sound inferences from such experiments relies on appropriate mathematical methods to model the distribution of reads along the genome, which has been challenging due to the scale and nature of these data.\n\nResultsWe propose a new framework (SRSFseq) based on Square Root Slope Functions shape analysis to analyse Illumina sequencing data. In the new approach the basic unit of information is the density of mapped reads over region of interest located on the known reference genome. The densities are interpreted as shapes and a new shape analysis model is proposed. An equivalent of a Fisher test is used to quantify the significance of shape differences in read distribution patterns between groups of density functions in different experimental conditions. We evaluated the performance of this new framework to analyze RNA-seq data at the exon level, which enabled the detection of variation in read distributions and abundances between experimental conditions not detected by other methods. Thus, the method is a suitable supplement to the state of the are count based techniques. The variety of density representations and flexibility of mathematical design allow the model to be easily adapted to other data types or problems in which the distribution of reads is to be tested. The functional interpretation and SRSF phase-amplitude separation technique gives an efficient noise reduction procedure improving the sensitivity and specificity of the method.

bioinformatics

Repli-seq: genome-wide analysis of replication timing by next-generation sequencing

Cycling cells duplicate their DNA content during S phase, following a defined program called replication timing (RT). Early and late replicating regions differ in terms of mutation rates, transcriptional activity, chromatin marks and sub-nuclear position. Moreover, RT is regulated during development and is altered in disease. Exploring mechanisms linking RT to other cellular processes in normal and diseased cells will be facilitated by rapid and robust methods with which to measure RT genome wide. Here, we describe a protocol to analyse genome-wide RT by next-generation sequencing (NGS). This protocol yields highly reproducible results across laboratories and platforms. We also provide the computational pipelines for analysis, parsing phased genomes using single nucleotide polymorphisms (SNP) for analyzing imprinted RT, and for direct comparison to Repli-chip data obtained by analyzing nascent DNA by microarrays.

genomics