bioRxiv ScienceSearch

Biology subjects

Noble, W. S.

Publications and source records attributed to Noble, W. S..

At least 19 recordsLinked to original sources

MoMo: Discovery of statistically significant post-translational modification motifs

MotivationPost-translational modifications (PTMs) of proteins are associated with many significant biological functions and can be identified in high throughput using tandem mass spectrometry. Many PTMs are associated with short sequence patterns called \"motifs\" that help localize the modifying enzyme. Accordingly, many algorithms have been designed to identify these motifs from mass spectrometry data. Accurate statistical confidence estimates for discovered motifs are critically important for proper interpretation and in the design of downstream experimental validation.\n\nResultsWe describe a method for assigning statistical confidence estimates to PTM motifs, and we demonstrate that this method provides accurate p-values on both simulated and real data. Our methods are implemented in MoMo, a software tool for discovering motifs among sets of PTMs that we make available as a web server and as downloadable source code. MoMo reimplements the two most widely used PTM motif discovery algorithms--motif-x and MoDL--while offering many enhancements. Relative to motif-x, MoMo offers improved statistical confidence estimates and more accurate calculation of motif scores. The MoMo web server offers more proteome databases, more input formats, larger inputs and longer running times than the motif-x web server. Finally, our study demonstrates that the confidence estimates produced by motif-x are inaccurate. This inaccuracy stems in part from the common practice of drawing \"background\" peptides from an unshuffled proteome database. Our results thus suggest that many of the hundreds of papers that use motif-x to find motifs may be reporting results that lack statistical support.\n\nAvailabilityhttp://meme-suite.org\n\nContacttimothybailey@unr.edu

bioinformatics

Multi-scale deep tensor factorization learns a latent representation of the human epigenome

The human epigenome has been experimentally characterized by measurements of protein binding, chromatin acessibility, methylation, and histone modification in hundreds of cell types. The result is a huge compendium of data, consisting of thousands of measurements for every basepair in the human genome. These data are difficult to make sense of, not only for humans, but also for computational methods that aim to detect genes and other functional elements, predict gene expression, characterize polymorphisms, etc. To address this challenge, we propose a deep neural network tensor factorization method, Avocado, that compresses epigenomic data into a dense, information-rich representation of the human genome. We use data from the Roadmap Epigenomics Consortium to demonstrate that this learned representation of the genome is broadly useful: first, by imputing epigenomic data more accurately than previous methods, and second, by showing that machine learning models that exploit this representation outperform those trained directly on epigenomic data on a variety of genomics tasks. These tasks include predicting gene expression, promoter-enhancer interactions, replication timing, and an element of 3D chromatin architecture. Our findings suggest the broad utility of Avocados learned latent representation for computational genomics and epigenomics.

bioinformatics

High-throughput mapping of meiotic crossover and chromosome mis-segregation events in interspecific hybrid mice

We developed \"sci-LIANTI\", a high-throughput, high-coverage single-cell DNA sequencing method that combines single-cell combinatorial indexing (\"sci\") and linear amplification via transposon insertion (\"LIANTI\"). To characterize rare chromosome mis-segregation events in male meiosis and their relationship to the landscape of meiotic crossovers, we applied sci-LIANTI to profile the genomes of 6,928 sperm and sperm precursors from infertile, interspecific F1 male mice. From 1,663 haploid and 292 diploid cells, we mapped 24,672 crossover events and identified genomic and epigenomic contexts that influence crossover hotness. Surprisingly, we observed frequent mitotic chromosome segregation during meiosis. Moreover, segregation during meiosis in individual cells was highly biased towards either mitotic or meiotic events. We anticipate that sci-LIANTI can be applied to fully characterize various recombination landscapes, as well as to other fields requiring high-throughput, high-coverage single-cell genome sequencing.\n\nOne Sentence SummarySingle-cell genome sequencing maps crossover and non-meiotic chromosome segregation during spermatogenesis in interspecific hybrid mice.

genetics

Joint precursor elution profile inference via regression for peptide detection in data-independent acquisition mass spectra

In data independent acquisition (DIA) mass spectrometry, precursor scans are interleaved with wide-window fragmentation scans, resulting in complex fragmentation spectra containing multiple co-eluting peptide species. In this setting, detecting the isotope distribution profiles of intact peptides in the precursor scans can be a critical initial step in accurate peptide detection and quantification. This peak detection step is particularly challenging when the isotope peaks associated with two different peptide species overlap--or interfere--with one another. We propose a regression model, called Siren, to detect isotopic peaks in precursor DIA data that can explicitly account for interference. We validate Sirens peak-calling performance on a variety of data sets by counting how many of the peaks Siren identifies are associated with confidently detected peptides. In particular, we demonstrate that substituting the Siren regression model in place of the existing peak-calling step in DIA-Umpire leads to improved overall rates of peptide detection.

bioinformatics

Fast open modification spectral library searching through approximate nearest neighbor indexing

Open modification searching (OMS) is a powerful search strategy that identifies peptides carrying any type of modification by allowing a modified spectrum to match against its unmodified variant by using a very wide precursor mass window. A drawback of this strategy, however, is that it leads to a large increase in search time. Although performing an open search can be done using existing spectral library search engines by simply setting a wide precursor mass window, none of these tools have been optimized for OMS, leading to excessive runtimes and suboptimal identification results. Here we present the ANN-SoLo tool for fast and accurate open spectral library searching. ANN-SoLo uses approximate nearest neighbor indexing to speed up OMS by selecting only a limited number of the most relevant library spectra to compare to an unknown query spectrum. This approach is combined with a cascade search strategy to maximize the number of identified unmodified and modified spectra while strictly controlling the false discovery rate, as well as a shifted dot product score to sensitively match modified spectra to their unmodified counterparts. ANN-SoLo achieves state-of-the-art performance in terms of speed and the number of identifications. On a previously published human cell line data set, ANN-SoLo confidently identifies more spectra than SpectraST or MSFragger and achieves a speedup of an order of magnitude compared to SpectraST.\n\nANN-SoLo is implemented in Python and C++. It is freely available under the Apache 2.0 license at https://github.com/bittremieux/ANN-SoLo.

bioinformatics

Cohesin interacts with a panoply of splicing factors required for cell cycle progression and genomic organization

The cohesin complex regulates sister chromatid cohesion, chromosome organization, gene expression, and DNA repair. Here we report that endogenous human cohesin interacts with a panoply of splicing factors and RNA binding proteins, including diverse components of the U4/U6.U5 tri-snRNP complex and several splicing factors that are commonly mutated in cancer. The interactions are enhanced during mitosis, and the interacting splicing factors and RNA binding proteins follow the cohesin cycle and prophase pathway of regulated interactions with chromatin. Depletion of cohesin-interacting splicing factors results in stereotyped cell cycle arrests and alterations in genomic organization. These data support the hypothesis that splicing factors and RNA binding proteins control cell cycle progression and genomic organization via regulated interactions with cohesin and chromatin.\n\nOne Sentence SummaryEndogenous tagging reveals that cohesin interacts with diverse chromatin-bound splicing factors that regulate cell cycle progression and genomic organization in human cells.

cell biology

Dynamic reorganization of nuclear architecture during human cardiogenesis

While chromosomal architecture varies among cell types, little is known about how this organization is established or its role in development. We integrated Hi-C, RNA-seq and ATAC-seq during cardiac differentiation from human pluripotent stem cells to generate a comprehensive profile of chromosomal architecture. We identified active and repressive domains that are dynamic during cardiogenesis and recapitulate in vivo cardiomyocytes. During differentiation, heterochromatic regions condense in cis. In contrast, many cardiac-specific genes, such as TTN (titin), decompact and transition to an active compartment coincident with upregulation. Moreover, we identify a network of genes, including TTN, that share the heart-specific splicing factor, RBM20, and become associated in trans during differentiation, suggesting the existence of a 3D nuclear splicing factory. Our results demonstrate both the dynamic nature in nuclear architecture and provide insights into how developmental genes are coordinately regulated.\n\nOne Sentence SummaryThe three-dimensional structure of the human genome is dynamically regulated both globally and locally during cardiogenesis.

genomics

Measuring the reproducibility and quality of Hi-C data

Hi-C is currently the most widely used assay to investigate the 3D organization of the genome and to study its role in gene regulation, DNA replication, and disease. However, Hi-C experiments are costly to perform and involve multiple complex experimental steps; thus, accurate methods for measuring the quality and reproducibility of Hi-C data are essential to determine whether the output should be used further in a study. Using real and simulated data, we profile the performance of several recently proposed methods for assessing reproducibility of population Hi-C data, including HiCRep, GenomeDISCO, HiC-Spector and QuASAR-Rep. By explicitly controlling noise and sparsity through simulations, we demonstrate the deficiencies of performing simple correlation analysis on pairs of matrices, and we show that methods developed specifically for Hi-C data produce better measures of reproducibility. We also show how to use established (e.g., ratio of intra to interchromosomal interactions) and novel (e.g., QuASAR-QC) measures to identify low quality experiments. In this work, we assess reproducibility and quality measures by varying sequencing depth, resolution and noise levels in Hi-C data from 13 cell lines, with two biological replicates each, as well as 176 simulated matrices. Through this extensive validation and benchmarking of Hi-C data, we describe best practices for reproducibility and quality assessment of Hi-C experiments. We make all software publicly available at http://github.com/kundajelab/3DChromatin_ReplicateQC to facilitate adoption in the community.

genomics

Using DNase Hi-C techniques to map global and local three-dimensional genome architecture at high resolution

The folding and three-dimensional (3D) organization of chromatin in the nucleus critically impacts genome function. The past decade has witnessed rapid advances in genomic tools for delineating 3D genome architecture. Among them, chromosome conformation capture (3C)-based methods such as Hi-C are the most widely used techniques for mapping chromatin interactions. However, traditional Hi-C protocols rely on restriction enzymes (REs) to fragment chromatin and are therefore limited in resolution. We recently developed DNase Hi-C for mapping 3D genome organization, which uses DNase I for chromatin fragmentation. DNase Hi-C overcomes RE-related limitations associated with traditional Hi-C methods, leading to improved methodological resolution. Furthermore, combining this method with DNA capture technology provides a high-throughput approach (targeted DNase Hi-C) that allows for mapping fine-scale chromatin architecture at exceptionally high resolution. Hence, targeted DNase Hi-C will be valuable for delineating the physical landscapes of cis-regulatory networks that control gene expression and for characterizing phenotype-associated chromatin 3D signatures. Here, we provide a detailed description of method design and step-by-step working protocols for these two methods.\n\nHighlightsO_LIDNase Hi-C, a method for comprehensive mapping of chromatin contacts on a whole-genome scale, is based on random chromatin fragmentation by DNase I digestion instead of sequence-specific restriction enzyme (RE) digestion.\nC_LIO_LITargeted DNase Hi-C, which combines DNase Hi-C with DNA capture technology, is a high-throughput method for mapping fine-scale chromatin architecture of genomic loci of interest at a resolution comparable to that of genomic annotations of functional elements.\nC_LIO_LIDNase Hi-C and targeted DNase Hi-C provide the first high-throughput way to overcome the RE-digestion-associated resolution limit of 3C-based methods.\nC_LIO_LIStep-by-step whole-genome and targeted DNase Hi-C protocols for mapping global and local 3D genome architecture, respectively, are described.\nC_LI

genomics

GenomeDISCO: A concordance score for chromosome conformation capture experiments using random walks on contact map graphs

MotivationThe three-dimensional organization of chromatin plays a critical role in gene regulation and disease. High-throughput chromosome conformation capture experiments such as Hi-C are used to obtain genome-wide maps of 3D chromatin contacts. However, robust estimation of data quality and systematic comparison of these contact maps is challenging due to the multi-scale, hierarchical structure of chromatin contacts and the resulting properties of experimental noise in the data. Measuring concordance of contact maps is important for assessing reproducibility of replicate experiments and for modeling variation between different cellular contexts.\n\nResultsWe introduce a concordance measure called GenomeDISCO (DIfferences between Smoothed COntact maps) for assessing the similarity of a pair of contact maps obtained from chromosome conformation capture experiments. The key idea is to smooth contact maps using random walks on the contact map graph, before estimating concordance. We use simulated datasets to benchmark GenomeDISCOs sensitivity to different types of noise that affect chromatin contact maps. When applied to a large collection of Hi-C datasets, GenomeDISCO accurately distinguishes biological replicates from samples obtained from different cell types. GenomeDISCO also generalizes to other chromosome conformation capture assays, such as HiChIP.\n\nAvailabilitySoftware implementing GenomeDISCO is available at https://github.com/kundajelab/genomedisco.\n\nContactakundaje@stanford.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Orientation-dependent Dxz4 contacts shape the 3D structure of the inactive X chromosome

The mammalian inactive X chromosome (Xi) condenses into a bipartite structure with two superdomains of frequent long-range contacts separated by a boundary or hinge region. Using in situ DNase Hi-C in mouse cells with deletions or inversions within the hinge we show that the conserved repeat locus Dxz4 alone is sufficient to maintain the bipartite structure and that Dxz4 orientation controls the distribution of long-range contacts on the Xi. Frequent long-range contacts between Dxz4 and the telomeric superdomain are either lost after its deletion or shifted to the centromeric superdomain after its inversion. This massive reversal in contact distribution is consistent with the reversal of CTCF motif orientation at Dxz4. De-condensation of the Xi after Dxz4 deletion is associated with partial restoration of TADs normally attenuated on the Xi. There is also an increase in chromatin accessibility and CTCF binding on the Xi after Dxz4 deletion or inversion, but few changes in gene expression, in accordance with multiple epigenetic mechanisms ensuring X silencing. We propose that Dxz4 represents a structural platform for frequent long-range contacts with multiple loci in a direction dictated by the orientation of a bank of CTCF motifs at Dxz4, which may work as a ratchet to form the distinctive bipartite structure of the condensed Xi.

molecular biology

Segway 2.0: Gaussian mixture models and minibatch training

SummarySegway performs semi-automated genome annotation, discovering joint patterns across multiple genomic signal datasets. We discuss a major new version of Segway and highlight its ability to model data with substantially greater accuracy. Major enhancements in Segway 2.0 include the ability to model data with a mixture of Gaussians, enabling capture of arbitrarily complex signal distributions, and minibatch training, leading to better learned parameters.\n\nAvailability and ImplementationSegway and its source code are freely available for download at https://segway.hoffmanlab.org. We have made available scripts (https://doi.org/10.5281/zenodo.802940) and datasets (https://doi.org/10.5281/zenodo.802907) for this papers analysis.\n\nContactmichael.hoffman@utoronto.ca

bioinformatics

PREDICTD: PaRallel Epigenomics Data Imputation With Cloud-based Tensor Decomposition

The Encyclopedia of DNA Elements (ENCODE) and the Roadmap Epigenomics Project have produced thousands of data sets mapping the epigenome in hundreds of cell types. However, the number of cell types remains too great to comprehensively map given current time and financial constraints. We present a method, PaRallel Epigenomics Data Imputation with Cloud-based Tensor Decomposition (PREDICTD), to address this issue by computationally imputing missing experiments in collections of epigenomics experiments. PREDICTD leverages an intuitive and natural model called \"tensor decomposition\" to impute many experiments simultaneously. Compared with the current state-of-the-art method, ChromImpute, PREDICTD produces lower overall mean squared error, and combining methods yields further improvement. We show that PREDICTD data can be used to investigate enhancer biology at non-coding human accelerated regions. PREDICTD provides reference imputed data sets and open-source software for investigating new cell types, and demonstrates the utility of tensor decomposition and cloud computing, two technologies increasingly applicable in bioinformatics.

bioinformatics

An Integrative Framework For Detecting Structural Variations In Cancer Genomes

Structural variants can contribute to oncogenesis through a variety of mechanisms, yet, despite their importance, the identification of structural variants in cancer genomes remains challenging. Here, we present an integrative framework for comprehensively identifying structural variation in cancer genomes. For the first time, we apply next-generation optical mapping, high-throughput chromosome conformation capture (Hi-C), and whole genome sequencing to systematically detect SVs in a variety of cancer cells.\n\nUsing this approach, we identify and characterize structural variants in up to 29 commonly used normal and cancer cell lines. We find that each method has unique strengths in identifying different classes of structural variants and at different scales, suggesting that integrative approaches are likely the only way to comprehensively identify structural variants in the genome. Studying the impact of the structural variants in cancer cell lines, we identify widespread structural variation events affecting the functions of non-coding sequences in the genome, including the deletion of distal regulatory sequences, alteration of DNA replication timing, and the creation of novel 3D chromatin structural domains.\n\nThese results underscore the importance of comprehensive structural variant identification and indicate that non-coding structural variation may be an underappreciated mutational process in cancer genomes.

genomics

HiCRep: assessing the reproducibility of Hi-C data using a stratum-adjusted correlation coefficient

Hi-C is a powerful technology for studying genome-wide chromatin interactions. However, current methods for assessing Hi-C data reproducibility can produce misleading results because they ignore spatial features in Hi-C data, such as domain structure and distance dependence. We present HiCRep, a framework for assessing the reproducibility of Hi-C data that systematically accounts for these features. In particular, we introduce a novel similarity measure, the stratum adjusted correlation coefficient (SCC), for quantifying the similarity between Hi-C interaction matrices. Not only does it provide a statistically sound and reliable evaluation of reproducibility, SCC can also be used to quantify differences between Hi-C contact matrices and to determine the optimal sequencing depth for a desired resolution. The measure consistently shows higher accuracy than existing approaches in distinguishing subtle differences in reproducibility and depicting interrelationships of cell lineages. The proposed measure is straightforward to interpret and easy to compute, making it well-suited for providing standardized, interpretable, automatable, and scalable quality control. The freely available R package HiCRep implements our approach.

bioinformatics

The dynamic three-dimensional organization of the diploid yeast genome

The budding yeast Saccharomyces cerevisiae is a long-standing model for the three-dimensional organization of eukaryotic genomes. Even in this well-studied model, it is unclear how homolog pairing in diploids and environment-induced gene relocalization influence overall genome organization. Here, we performed high-throughput chromosome conformation capture on diverged Saccharomyces hybrid diploids to obtain the first global view of chromosome conformation in diploid yeasts. After controlling for the Rabl-like orientation, we observe significant homolog proximity that increased in saturated culture conditions. Surprisingly, we observe a localized increase in homologous interactions between the HAS1 alleles specifically under galactose induction and saturated growth, mediated by association with nuclear pore complexes at the nuclear periphery. Together, these results reveal that the diploid yeast genome has a dynamic and complex 3D organization.

genomics

DNA sequence+shape kernel enables alignment-free modeling of transcription factor binding

MotivationTranscription factors (TFs) bind to specific DNA sequence motifs. Several lines of evidence suggest that TF-DNA binding is mediated in part by properties of the local DNA shape: the width of the minor groove, the relative orientations of adjacent base pairs, etc. Several methods have been developed to jointly account for DNA sequence and shape properties in predicting TF binding affinity. However, a limitation of these methods is that they typically require a training set of aligned TF binding sites.\n\nResultsWe describe a sequence+shape kernel that leverages DNA sequence and shape information to better understand protein-DNA binding preference and affinity. This kernel extends an existing class of k-mer based sequence kernels, based on the recently described di-mismatch kernel. Using three in vitro benchmark datasets, derived from universal protein binding microarrays (uPBMs), genomic context PBMs (gcPBMs) and SELEX-seq data, we demonstrate that incorporating DNA shape information improves our ability to predict protein-DNA binding affinity. In particular, we observe that (1) the k-spectrum+shape model performs better than the classical k-spectrum kernel, particularly for small k values; (2) the di-mismatch kernel performs better than the k-mer kernel, for larger k; and (3) the di-mismatch+shape kernel performs better than the di-mismatch kernel for intermediate k values.\n\nAvailabilityThe software is available at https://bitbucket.org/wenxiu/sequence-shape.git\n\nContactrohs@usc.edu, william-noble@uw.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

HiC-Spector: A matrix library for spectral and reproducibility analysis of Hi-C contact maps

SummaryGenome-wide proximity ligation based assays like Hi-C have opened a window to the 3D organization of the genome. In so doing, they present data structures that are different from conventional 1D signal tracks. To exploit the 2D nature of Hi-C contact maps, matrix techniques like spectral analysis are particularly useful. Here, we present HiC-spector, a collection of matrix-related functions for analyzing Hi-C contact maps. In particular, we introduce a novel reproducibility metric for quantifying the similarity between contact maps based on spectral decomposition. The metric successfully separates contact maps mapped from Hi-C data coming from biological replicates, pseudo-replicates and different cell types.\n\nAvailabilitySource code in Julia and the documentation of HiC-spector can be freely obtained at https://github.com/gersteinlab/HiC_spector\n\nContactpi@gersteinlab.org

bioinformatics