bioRxiv Science⌕ Search

Biology subjects

Pani, S.

Publications and source records attributed to Pani, S..

2 recordsLinked to original sources

Long-read sequencing and structural variant characterization in 1,019 samples from the 1000 Genomes Project

Structural variants (SVs) contribute significantly to human genetic diversity and disease1-4. Previously, SVs have remained incompletely resolved by population genomics, with short-read sequencing facing limitations in capturing the whole spectrum of SVs at nucleotide resolution5-7. Here we leveraged nanopore sequencing8 to construct an intermediate coverage resource of 1,019 long-read genomes sampled within 26 human populations from the 1000 Genomes Project. By integrating linear and graph-based approaches for SV analysis via pangenome graph-augmentation, we uncover 167,291 sequence-resolved SVs in these samples, considerably advancing SV characterization compared to population-wide short-read sequencing studies3,4. Our analysis details diverse SV classes--deletions, duplications, insertions, and inversions--at population-scale. LINE-1 and SVA retrotransposition activities frequently mediate transductions9,10 of unique sequences, with both mobile element classes transducing sequences at either the 3'- or 5'-end, depending on the source element locus. Furthermore, analyses of SV breakpoint junctions suggest a continuum of homology-mediated rearrangement processes are integral to SV formation, and highlight evidence for SV recurrence involving repeat sequences. Our open-access dataset underscores the transformative impact of long-read sequencing in advancing the characterisation of polymorphic genomic architectures, and provides a resource for guiding variant prioritisation in future long-read sequencing-based disease studies.

genomics↗

SVarp: pangenome-based structural variant discovery

The linear human reference genome that we use today does not represent the haplotypic diversity of the global human population. This raises bias in genomic read alignment and limits our ability to call large structural variations (SV), especially at highly polymorphic loci. Thus, many SV alleles remain unresolved. Recent efforts to transition to a graph-based reference genome resulted in the generation of the first draft human pangenome reference, but tools to call SVs relative to the pangenome reference are presently lacking. In this study, we present the SVarp algorithm, aiming to discover haplotype resolved SVs on top of a pangenome reference using long sequencing reads. SVarp outputs local assemblies of SV alleles, termed svtigs, instead of a VCF file of SV breakpoints, which we propose as a general exchange format allowing for flexible downstream analyses. In order to assess the accuracy of svtigs, we used simulated and real human genomes. Simulations allowed us to make exact breakpoint comparisons against the true callsets. We observed [~]96% recall with deletions, insertions and duplications larger than 1,000bp, showing that SVarp can reliably detect genomic structural variants not yet represented in the graph. On the other hand, we compared SVarp output for ONT sequencing data at 20X coverage against independent genome assemblies of the same samples and found that [~]82% of our svtig predictions are validated by the assemblies by a match with more than 85% sequence identity. SVarp was implemented using C++ and its source code is available at https://github.com/asylvz/SVarp under MIT license.

bioinformatics↗