bioRxiv Science⌕ Search

Biology subjects

Ranchalis, J.

Publications and source records attributed to Ranchalis, J..

5 recordsLinked to original sources

A haplotype-resolved view of human gene regulation

Diploid human cells contain two non-identical genomes, and differences in their regulation underlie human development and disease. We present Fiber-seq Inferred Regulatory Elements (FIRE) and show that FIRE provides a more comprehensive and quantitative snapshot of the accessible chromatin landscape across the 6 Gbp diploid human genome, overcoming previously unrecognized biases in existing regulatory element catalogs. FIRE enables comprehensive detection of haplotype-selective chromatin accessibility (HSCA), exposing novel imprinted elements lacking underlying parent-of-origin CpG methylation differences, and gene regulatory modules that permit genes to escape X chromosome inactivation. We uncover that the human leukocyte antigen (HLA) locus harbors the most HSCA in immune cells, where we resolve specific transcription factor (TF) binding events disrupted by disease-associated variants. Finally, we demonstrate that the regulatory landscape of a cell is littered with autosomal somatic chromatin epimutations that are propagated by clonal expansions to create mitotically stable and non-genetically deterministic chromatin alterations.

genomics↗

RNA polymerases reshape chromatin and coordinate transcription on individual fibers

During eukaryotic transcription, RNA polymerases must initiate and pause within a crowded, complex environment, surrounded by nucleosomes and other transcriptional activity. This environment creates a spatial arrangement along individual chromatin fibers ripe for both competition and coordination, yet these relationships remain largely unknown owing to the inherent limitations of traditional structural and sequencing methodologies. To address these limitations, we employed long-read chromatin fiber sequencing (Fiber-seq) to visualize RNA polymerases within their native chromatin context at single-molecule and near single-nucleotide resolution along up to 30 kb fibers. We demonstrate that Fiber-seq enables the identification of single-molecule RNA Polymerase (Pol) II and III transcription associated foot-prints, which, in aggregate, mirror bulk short-read sequencing-based measurements of transcription. We show that Pol II pausing destabilizes downstream nucleosomes, with frequently paused genes maintaining a short-term memory of these destabilized nucleosomes. Furthermore, we demonstrate pervasive direct coordination and anti-coordination between nearby Pol II genes, Pol III genes, transcribed enhancers, and insulator elements. This coordination is largely limited to spatially organized elements within 5 kb of each other, implicating short-range chromatin environments as a predominant determinant of coordinated polymerase initiation. Overall, transcription initiation reshapes surrounding nucleosome architecture and coordinates nearby transcriptional machinery along individual chromatin fibers.

genomics↗

Synchronized long-read genome, methylome, epigenome, and transcriptome for resolving a Mendelian condition

Resolving the molecular basis of a Mendelian condition (MC) remains challenging owing to the diverse mechanisms by which genetic variants cause disease. To address this, we developed a synchronized long-read genome, methylome, epigenome, and transcriptome sequencing approach, which enables accurate single-nucleotide, insertion-deletion, and structural variant calling and diploid de novo genome assembly, and permits the simultaneous elucidation of haplotype-resolved CpG methylation, chromatin accessibility, and full-length transcript information in a single long-read sequencing run. Application of this approach to an Undiagnosed Diseases Network (UDN) participant with a chromosome X;13 balanced translocation of uncertain significance revealed that this translocation disrupted the functioning of four separate genes (NBEA, PDK3, MAB21L1, and RB1) previously associated with single-gene MCs. Notably, the function of each gene was disrupted via a distinct mechanism that required integration of the four omes to resolve. These included nonsense-mediated decay, fusion transcript formation, enhancer adoption, transcriptional readthrough silencing, and inappropriate X chromosome inactivation of autosomal genes. Overall, this highlights the utility of synchronized long-read multi-omic profiling for mechanistically resolving complex phenotypes.

genetics↗

Fibertools: fast and accurate DNA-m6A calling using single-molecule long-read sequencing

Long-read DNA sequencing has recently emerged as a powerful tool for studying both genetic and epigenetic architectures at single-molecule and single-nucleotide resolution. Long-read epigenetic studies encompass both the direct identification of native cytosine methylation as well as the identification of exogenously placed DNA N6-methyladenine (DNA-m6A). However, detecting DNA-m6A modifications using single-molecule sequencing, as well as co-processing single-molecule genetic and epigenetic architectures, is limited by computational demands and a lack of supporting tools. Here, we introduce fibertools, a state-of-the-art toolkit that features a semi-supervised convolutional neural network for fast and accurate identification of m6A-marked bases using PacBio single-molecule long-read sequencing, as well as the co-processing of long-read genetic and epigenetic data produced using either PacBio or Oxford Nanopore sequencing platforms. We demonstrate accurate DNA-m6A identification (>90% precision and recall) along >20 kilobase long DNA molecules with a [~]1,000-fold improvement in speed. In addition, we demonstrate that fibertools can readily integrate genetic and epigenetic data at single-molecule resolution, including the seamless conversion between molecular and reference coordinate systems, allowing for accurate genetic and epigenetic analyses of long-read data within structurally and somatically variable genomic regions.

bioinformatics↗

Full-length isoform sequencing for resolving the molecular basis of Charcot-Marie-Tooth 2A

ObjectivesTranscript sequencing of patient derived samples has been shown to improve the diagnostic yield for solving cases of likely Mendelian disorders, yet the added benefit of full-length long-read transcript sequencing is largely unexplored. MethodsWe applied short-read and full-length isoform cDNA sequencing and mitochondrial functional studies to a patient-derived fibroblast cell line from an individual with neuropathy that previously lacked a molecular diagnosis. ResultsWe identified an intronic homozygous MFN2 c.600-31T>G variant that disrupts a branch point critical for intron 6 spicing. Full-length long-read isoform cDNA sequencing after treatment with a nonsense-mediated mRNA decay (NMD) inhibitor revealed that this variant creates five distinct altered splicing transcripts. All five altered splicing transcripts have disrupted open reading frames and are subject to NMD. Furthermore, a patient-derived fibroblast line demonstrated abnormal lipid droplet formation, consistent with MFN2 dysfunction. Although correctly spliced full-length MFN2 transcripts are still produced, this branch point variant results in deficient MFN2 protein levels and autosomal recessive Charcot-Marie-Tooth disease, axonal, type 2A (CMT2A). DiscussionThis case highlights the utility of full-length isoform sequencing for characterizing the molecular mechanism of undiagnosed rare diseases and expands our understanding of the genetic basis for CMT2A.

genetics↗