bioRxiv Science⌕ Search

Biology subjects

Beck, C. R.

Publications and source records attributed to Beck, C. R..

6 recordsLinked to original sources

Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall

Advances in long-read sequencing (LRS) technology continue to make whole-genome sequencing more complete, affordable, and accurate. LRS provides significant advantages over short-read sequencing approaches, including phased de novo genome assembly, access to previously excluded genomic regions, and discovery of more complex structural variants (SVs) associated with disease. Limitations remain with respect to cost, scalability, and platform-dependent read accuracy and the tradeoffs between sequence coverage and sensitivity of variant discovery are important experimental considerations for the application of LRS. We compare the genetic variant calling precision and recall of Oxford Nanopore Technologies (ONT) and PacBio HiFi platforms over a range of sequence coverages. For read-based applications, LRS sensitivity begins to plateau around 12-fold coverage with a majority of variants called with reasonable accuracy (F1 score above 0.5), and both platforms perform well for SV detection. Genome assembly increases variant calling precision and recall of SVs and indels in HiFi datasets with HiFi outperforming ONT in quality as measured by the F1 score of assembly-based variant callsets. While both technologies continue to evolve, our work offers guidance to design cost-effective experimental strategies that do not compromise on discovering novel biology.

genomics↗

Assembly of 43 diverse human Y chromosomes reveals extensive complexity and variation

The prevalence of highly repetitive sequences within the human Y chromosome has led to its incomplete assembly and systematic omission from genomic analyses. Here, we present long-read de novo assemblies of 43 diverse Y chromosomes spanning 180,000 years of human evolution, including two from deep-rooted African Y lineages, and report remarkable complexity and diversity in chromosome size and structure, in contrast with its low level of base substitution variation. The size of the Y chromosome assemblies varies extensively from 45.2 to 84.9 Mbp and include, on average, 81 kbp of novel sequence per Y chromosome. Half of the male-specific euchromatic region is subject to large inversions with a >2-fold higher recurrence rate compared to inversions in the rest of the human genome. Ampliconic sequences associated with these inversions further show differing mutation rates that are sequence context-dependent and some ampliconic genes show evidence for concerted evolution with the acquisition and purging of lineage-specific pseudogenes. The largest heterochromatic region in the human genome, the Yq12, is composed of alternating arrays of DYZ1 and DYZ2 repeat units that show extensive variation in the number, size and distribution of these arrays, but retain a 1:1 copy number ratio of the monomer repeats, consistent with the notion that functional or evolutionary forces are acting on this chromosomal region. Finally, our data suggests that the boundary between the recombining pseudoautosomal region 1 and the non-recombining portions of the X and Y chromosomes lies 500 kbp distal to the currently established boundary. The availability of sequence-resolved Y chromosomes from multiple individuals provides a unique opportunity for identifying new associations of specific traits with Y-chromosomal variants and garnering novel insights into the evolution and function of complex regions of the human genome.

genomics↗

Resolution of structural variation in diverse mouse genomes reveals chromatin remodeling due to transposable elements

Diverse inbred mouse strains are among the foremost models for biomedical research, yet genome characterization of many strains has been fundamentally lacking in comparison to human genomics research. In particular, the discovery and cataloging of structural variants is incomplete, limiting the discovery of potentially causative alleles for phenotypic variation across individuals. Here, we utilized long-read sequencing to resolve genome-wide structural variants (SVs, variants [≥] 50 bp) in 20 genetically distinct inbred mice. We report 413,758 site-specific SVs that affect 13% (356 Mbp) of the current mouse reference assembly, including 510 previously unannotated variants which alter coding sequences. We find that 39% of SVs are attributed to transposable element (TE) variation accounting for 75% of bases altered by SV. We then utilized this callset to investigate the impact of TE heterogeneity on mouse embryonic stem cells (mESCs), and find multiple TE classes that influence chromatin accessibility across loci. We also identify strain-specific transcription start sites originating in polymorphic TEs that modify gene expression. Our work provides the first long-read based analysis of mouse SVs and illustrates that previously unresolved TEs underlie epigenetic and transcriptome differences in mESCs.

genomics↗

Transposable element-mediated rearrangements are prevalent in human genomes

Transposable elements constitute about half of human genomes, and their role in generating human variation through retrotransposition is broadly studied and appreciated. Structural variants mediated by transposons, which we call transposable element-mediated rearrangements (TEMRs), are less well studied, and the mechanisms leading to their formation as well as their broader impact on human diversity are poorly understood. Here, we identify 493 unique TEMRs across the genomes of three individuals. While homology directed repair is the dominant driver of TEMRs, our sequence-resolved TEMR resource allows us to identify complex inversion breakpoints, triplications or other high copy number polymorphisms, and additional complexities. TEMRs are enriched in genic loci and can create potentially important risk alleles such as a deletion in TRIM65, a known cancer biomarker and therapeutic target. These findings expand our understanding of this important class of structural variation, the mechanisms responsible for their formation, and establish them as an important driver of human diversity.

genomics↗

Haplotype-resolved inversion landscape reveals hotspots of mutational recurrence associated with genomic disorders

Unlike copy number variants (CNVs), inversions remain an underexplored genetic variation class. By integrating multiple genomic technologies, we discover 729 inversions in 41 human genomes. Approximately 85% of inversions <2 kbp form by twin-priming during L1-retrotransposition; 80% of the larger inversions are balanced and affect twice as many base pairs as CNVs. Balanced inversions show an excess of common variants, and 72% are flanked by segmental duplications (SDs) or mobile elements. Since this suggests recurrence due to non-allelic homologous recombination, we developed complementary approaches to identify recurrent inversion formation. We describe 40 recurrent inversions encompassing 0.6% of the genome, showing inversion rates up to 2.7x10-4 per locus and generation. Recurrent inversions exhibit a sex- chromosomal bias, and significantly co-localize to the critical regions of genomic disorders. We propose that inversion recurrence results in an elevated number of heterozygous carriers and structural SD diversity, which increases mutability in the population and predisposes to disease- causing CNVs.

genomics↗

Long-read isoform sequencing reveals survival-associated splicing in breast cancer

Tumors display widespread transcriptome alterations, but the full repertoire of isoform-level alternative splicing in cancer is not known. We developed a long-read RNA sequencing and analytical platform that identifies and annotates full-length isoforms, and infers tumor-specific splicing events. Application of this platform to breast cancer samples vastly expands the known isoform landscape of breast cancer, identifying thousands of previously unannotated isoforms of which ~30% impact protein coding exons and are predicted to alter protein localization and function, including of the breast cancer-associated genes ESR1 and ERBB2. We performed extensive cross-validation with -omics data sets to support transcription and translation of novel isoforms. We identified 3,059 breast tumor-specific splicing events, including 35 that are significantly associated with patient survival. Together, our results demonstrate the complexity, cancer subtype-specificity, and clinical relevance of novel isoforms in breast cancer that are only annotatable by LR-seq, and provide a rich resource of immuno-oncology therapeutic targets.

genomics↗