bioRxiv Science⌕ Search

bioRxiv · 10.1101/2023.12.01.569684

HiCDiff: single-cell Hi-C data denoising with diffusion models

Abstract

The genome-wide single-cell chromosome conformation capture technique, i.e., single-cell Hi-C (ScHi-C), was recently developed to interrogate the conformation of the genome of individual cells. However, single-cell Hi-C data are much sparser and noisier than bulk Hi-C data of a population of cells, making it difficult to apply and analyze them in biological research. Here, we developed the first generative diffusion models (HiCDiff) to denoise single-cell Hi-C data in the form of chromosomal contact matrices. HiCDiff uses a deep residual network to remove the noise in the reverse process of diffusion and can be trained in both unsupervised and supervised learning modes. Benchmarked on several single-cell Hi-C test datasets, the diffusion models substantially remove the noise in single-cell Hi-C data. The unsupervised HiCDiff outperforms most supervised non-diffusion deep learning methods and achieves the performance comparable to the state-of-the-art supervised deep learning method in terms of multiple metrics, demonstrating that diffusion models are a useful approach to denoising single-cell Hi-C data. Moreover, its good performance holds on denoising bulk Hi-C data.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

wang, y., Cheng, J.. 2023-12-04. HiCDiff: single-cell Hi-C data denoising with diffusion models. https://doi.org/10.1101/2023.12.01.569684

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Uniformly processed transcriptome-wide alternative splicing profiles for pediatric cancer research

Alterations in regulatory processes like alternative splicing contribute to pediatric cancer development. Although splicing aberrations have been observed in pediatric leukemias, alternative splicing has yet to be studied in pediatric cancers at scale, due to a lack of uniformly processed, sample-level pediatric cancer splicing profiles with non-diseased tissue comparators. We address this need by quantifying splice event usage for a curated set of bulk RNA-seq datasets from the NCI's Therapeutically Applicable Research to Generate Effective Treatments (TARGET, n = 1152) and Genotype-Tissue Expression (GTEx, n = 1098) as a comparator. This Treehouse Splice Compendium is accompanied by a reproducible workflow that was used to generate the data in the compendium and reflects the largest known RNA-seq dataset processed by the splice quantification tool Shiba. The compendium is part of a suite of large, uniformly processed datasets aggregated by the UCSC Treehouse Childhood Cancer Initiative and Alex's Lemonade Stand Foundation's Childhood Cancer Data Lab, which include the Treehouse Expression Compendia, refine.bio, and the Single-cell Pediatric Cancer Atlas.

bioinformatics↗

Correlation-aware discovery of co-occurring mutational signatures in cancer

Somatic mutations in cancer genomes record the activities of diverse mutational processes. Mutational signature analysis has advanced mechanistic understanding of mutagenesis and informed clinical decision-making, yet existing methods assume independence among signatures--an unrealistic assumption that can produce composite or contaminated signatures, reduce detection power, and yield inconsistent results. Here we present Cornet (CORrelated NMF ExTraction), a framework for mutational signature discovery that explicitly models co-occurring processes and jointly infers signatures and their correlation structure. Benchmarking on simulated data shows Cornet more accurately recovers distinct signatures under strong correlations. Applied to cancer genomes, Cornet enables unsupervised discovery of the colibactin-associated signature SBS88 in oral cancers and identifies the tobacco smoking signature SBS4 in bladder cancer, where it was previously thought absent. Cornet also uncovers a novel mutational process implicated in early-onset colorectal cancer and a signature arising from the interplay between tobacco smoking and ERCC2-mutation-driven nucleotide-excision repair deficiency. Together, these results demonstrate that modeling correlations among mutational processes is essential for high-resolution signature discovery and dissecting the mutational etiology of human cancer.

bioinformatics↗

Individual-level expression deconvolution and assessment of cross-sample variation

Recovering cell-type-specific gene expression from bulk RNA sequencing would facilitate the study of transcriptional variation among individuals. However, accuracy can differ substantially among genes and cell types. We describe a reference-informed Bayesian deconvolution framework and a score that identifies gene--cell-type pairs likely to have more accurate estimates of cross-sample variation. The score uses bulk counts, reference expression profiles, and estimated RNA proportions. Known component expression is used to train and evaluate the score, but is not needed to calculate predictions from a trained model. We evaluated the approach in a ROSMAP-derived simulation with 40 target donors, 2,000 genes, and seven cell types. Median gene-wise correlation was 0.801 for raw allocated counts and 0.296 after normalization within each donor and cell type. To evaluate the score, we divided genes into five sets, kept linked genes together, and scored each set using a model trained on the other four. Retaining approximately 20\% of pairs within each cell type increased the median normalized correlation to 0.622. Ranking pairs only by the estimated share of a gene's bulk RNA contributed by the cell type yielded 0.570 at the same retained count. These results show that observable information can help prioritize pairs with more accurately recovered cross-sample variation.

bioinformatics↗