bioRxiv Science⌕ Search

Biology subjects

Gardner, J. M. V.

Publications and source records attributed to Gardner, J. M. V..

4 recordsLinked to original sources

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortiums (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

genomics↗

Haplotype-resolved centromeric chromatin organization from a complete diploid human genome

Centromeres ensure proper chromosome segregation during cell division, yet the organization and regulation of centromeric chromatin within satellite DNA arrays remain incompletely understood. Here, we leverage the complete diploid human genome benchmark (T2T-HG002) to provide a detailed study of centromeric sequence and chromatin architecture on individual haplotypes. Using adaptive-sampling-enriched, ultra-long-read DiMeLo-seq, we achieve single-molecule chromatin profiling across all centromeres, revealing that along single chromatin fibers, CENP-A, the histone variant specifying centromere identity, forms multiple discrete subdomains within hypomethylated centromere dip regions (CDRs) that are flanked by H3K9me3-enriched heterochromatin. Despite underlying sequence variation, CDRs localize to sequence-homogeneous domains and maintain relatively balanced CENP-A dosage and aggregate length across all chromosomes and between haplotypes. Further, we show that bidirectional changes to centromeric and pericentromeric DNA methylation are accompanied by changes to centromeric chromatin architecture. In passaged cells with centromeric hypomethylation, subdomain boundaries are eroded, and adjacent CENP-A domains tend to merge and expand. Conversely, in pluripotent stem cells with centromeric hypermethylation, CDRs are fundamentally reorganized, such that discrete hypomethylated domains are frequently consolidated into broader contiguous tracts. These methylation-associated CDR restructuring events suggest that DNA methylation acts as a principal regulator of human centromere organization, with implications for understanding centromere plasticity, epigenetic inheritance, and chromosomal instability in development and disease.

genomics↗

Complete genomes of a multi-generational pedigree to expand studies of genetic and epigenetic inheritance

Pedigree analysis remains the gold standard for rare disease diagnostics, yet whole genome sequencing studies typically omit critical regions like centromeres, telomeres, and acrocentric chromosome p-arms. Here, we present telomere-to-telomere (T2T) reference genomes for four self-identified African American individuals of admixed ancestry spanning three generations. Our parent-of-origin assigned, chromosome-level assemblies revealed precise meiotic recombination breakpoints in previously inaccessible regions, including recombination events across acrocentric and subtelomeric sequences. Centromeric regions were highly stable, with multi-megabase arrays inherited intact across three generations, while the position of kinetochore assembly sites remained consistent and predominantly associated with the p-arm proximal region. The relative lengths of telomeres on individual chromosomes were maintained across generations. Using a targeted rDNA assembly approach, we reconstructed a complete megabase-scale ribosomal DNA (rDNA) array corresponding to the paternal chromosome 14. This openly available pedigree provides a benchmark dataset for studying recombination and genetic and epigenetic variation across the complete genome.

genomics↗

RNA liquid biopsy via nanopore sequencing for novel biomarker discovery and cancer early detection

Liquid biopsies detect disease noninvasively by profiling cell-free nucleic acids that are secreted into the circulation. However, existing methods exhibit low sensitivity for detecting early stages of diseases such as cancer. Here we show that long-read nanopore sequencing of full-length cell-free RNA in plasma from healthy individuals, precancerous Barretts esophagus patients with high-grade dysplasia, or patients with esophageal adenocarcinoma reveals a diverse cell-free RNA transcriptome that can be leveraged for detecting and treating disease. We discovered 270,679 novel, intergenic cell-free RNAs, which we used to build a custom transcriptome reference for quantification, feature selection, and machine learning to accurately classify both precancer and cancer. Moreover, we found potential therapeutic targets, including metabolic, signaling, and immune checkpoint pathways, that are highly upregulated in both precancer and cancer patients. Our findings highlight the utility of our RNA liquid biopsy platform technology for discovering and targeting early stages of disease with molecular precision.

genomics↗