bioRxiv Science⌕ Search

Biology subjects

Mokveld, T.

Publications and source records attributed to Mokveld, T..

7 recordsLinked to original sources

Isocall enables scalable transcript identification from long-read RNA-sequencing data

Long-read RNA sequencing directly resolves the full structures of RNA transcripts. Advances in throughput now enable the generation of deeply sequenced cohorts of hundreds of samples, making joint transcript discovery across large datasets possible. However, existing transcript identification methods were designed for small datasets, which limits their applicability at this scale. Here, we present Isocall, a scalable and deterministic computational method for jointly calling transcripts from multiple PacBio long-read RNA sequencing samples. Isocall converts aligned full-length non-concatemer reads into compact per-sample transcript profiles, merges these profiles across samples, and jointly identifies known and novel transcripts supported by reads in the analyzed dataset. Filtering is tunable: presets provide coarse control and individual parameters, including relative abundance and internal priming thresholds, provide fine control. Isocall demonstrated high precision in our accuracy benchmarks, including the WTC11 SIRV spike-in controls, for which Isocall reported 0-2 false-positive transcripts per sample across the three SIRV mixes at default settings. To demonstrate scalability, we applied Isocall to 206 samples, totalling 3.5 billion raw reads, from the Human Pangenome Reference Consortium. After parallelized pbmm2 alignment and Isocall profile, the call step performed joint calling across the entire dataset in 25 minutes, using 1.3 GB of peak memory and 8 threads. Finally, in Genome in a Bottle samples with matched SNP genotypes, splice-site polymorphisms provide an additional measure of call accuracy: Isocall recovered 337 polymorphic splice sites, including a de novo donor site in BTN3A1 that corresponds to a complete isoform switch on the mutant allele.

bioinformatics↗

A computational model for quantifying instability of tandem repeats across the genome

Tandem repeats (TRs) exhibit high levels of somatic mosaicism, which is increasingly recognized as an important modifier of repeat expansion disorders. Long-read sequencing can capture full-length repeat alleles, yet robust frameworks for quantifying instability across TRs genome-wide are still needed. Here, we introduce a general-purpose model for quantifying TR instability in a given long-read sequencing dataset, without explicitly distinguishing biological mosaicism from technical noise, and which is broadly applicable to both simple and structurally complex loci. This model accurately characterizes allelic instability at each TR locus by representing the distribution of read-to-consensus deviations for each allele. Using HiFi sequencing data from 256 HPRC cell line samples, we fitted models for 617,007 TR loci, including known pathogenic repeats. We observe that instability levels are generally low, but vary substantially across individual TRs, and are driven more strongly by repeat composition than overall repeat length. Furthermore, we applied our method to targeted PureTarget long-read data from samples with known repeat expansions and identified significant mosaicism in the majority of expanded alleles. Our model offers a practical way to quantify instability of tandem repeats across the genome and to detect unusually unstable repeat alleles.

bioinformatics↗

A family portrait of the genomic factors shaping tandem repeat mutagenesis

Tandem repeats (TRs) are among the most mutable loci in the human genome, but the genomic determinants of TR mutagenesis remain mysterious. We used PacBio HiFi long-read sequencing to profile nearly eight million TR loci in 28 members of a large, four-generation CEPH/Utah family designated K1463. We identified 1,270 de novo TR expansions and contractions across 20 children in the pedigree. De novo mutations (DNMs) were more likely to occur at loci that were longer, composed of uninterrupted motif sequences, and heterozygous in the parental germline. Children born to older fathers also exhibited more de novo mutations at short tandem repeats (STRs). A total of 43 TR loci were hyper-mutable in K1463, expanding or contracting up to twelve times across the pedigree. Among hyper-mutable loci that comprised multiple motifs (i.e., "complex" loci), specific motifs expanded and contracted more often than others; for example, all ten DNMs at a complex, hyper-mutable locus near the non-coding RNA LINC03021 involved the same 19bp motif. The mutability of particular motifs may be attributable to allele length, as 95% of DNMs at complex loci were expansions and contractions of the most abundant motif on a parental haplotype. However, future work will be required to disentangle the effects of nucleotide content and allele length on motif-specific mutability, especially at hyper-mutable TRs. Overall, this study combines long-read sequencing technologies with new software tools to comprehensively investigate the factors that influence TR mutagenesis.

genomics↗

The Platinum Pedigree: A long-read benchmark for genetic variants

Recent advances in genome sequencing have improved variant calling in complex regions of the human genome. However, it is difficult to quantify variant calling performance since existing standards often focus on specificity, neglecting completeness in difficult to analyze regions. To create a more comprehensive truth set, we used Mendelian inheritance in a large pedigree (CEPH-1463) to filter variants across Illumina, PacBio high-fidelity (HiFi), and Oxford Nanopore Technologies platforms. This generated a variant map with over 4.7 million single-nucleotide variants, 767,795 indels, 537,486 tandem repeats, and 24,315 structural variants, covering 2.77 Gb of the GRCh38 genome. This work adds [~]200 Mb of high-confidence regions, including 8% more small variants, and introduces the first tandem repeat and structural variant truth sets for NA12878. As an example of the value of this improved benchmark, we retrained DeepVariant using this data to reduce genotyping errors by [~]34%.

genomics↗

A familial, telomere-to-telomere reference for human de novo mutation and recombination from a four-generation pedigree

Using five complementary short- and long-read sequencing technologies, we phased and assembled >95% of each diploid human genome in a four-generation, 28-member family (CEPH 1463) allowing us to systematically assess de novo mutations (DNMs) and recombination. From this family, we estimate an average of 192 DNMs per generation, including 75.5 de novo single-nucleotide variants (SNVs), 7.4 non-tandem repeat indels, 79.6 de novo indels or structural variants (SVs) originating from tandem repeats, 7.7 centromeric de novo SVs and SNVs, and 12.4 de novo Y chromosome events per generation. STRs and VNTRs are the most mutable with 32 loci exhibiting recurrent mutation through the generations. We accurately assemble 288 centromeres and six Y chromosomes across the generations, documenting de novo SVs, and demonstrate that the DNM rate varies by an order of magnitude depending on repeat content, length, and sequence identity. We show a strong paternal bias (75-81%) for all forms of germline DNM, yet we estimate that 17% of de novo SNVs are postzygotic in origin with no paternal bias. We place all this variation in the context of a high-resolution recombination map ([~]3.5 kbp breakpoint resolution). We observe a strong maternal recombination bias (1.36 maternal:paternal ratio) with a consistent reduction in the number of crossovers with increasing paternal (r=0.85) and maternal (r=0.65) age. However, we observe no correlation between meiotic crossover locations and de novo SVs, arguing against non-allelic homologous recombination as a predominant mechanism. The use of multiple orthogonal technologies, near-telomere-to-telomere phased genome assemblies, and a multi-generation family to assess transmission has created the most comprehensive, publicly available "truth set" of all classes of genomic variants. The resource can be used to test and benchmark new algorithms and technologies to understand the most fundamental processes underlying human genetic variation.

genomics↗

TRGT-denovo: accurate detection of de novo tandem repeat mutations

MotivationIdentifying de novo tandem repeat (TR) mutations on a genome-wide scale is essential for understanding genetic variability and its implications in rare diseases. While PacBio HiFi sequencing data enhances the accessibility of the genomes TR regions for genotyping, simple de novo calling strategies often generate an excess of likely false positives, which can obscure true positive findings, particularly as the number of surveyed genomic regions increases. ResultsWe developed TRGT-denovo, a computational method designed to accurately identify all types of de novo TR mutations--including expansions, contractions, and compositional changes-- within family trios. TRGT-denovo directly interrogates read evidence, allowing for the detection of subtle variations often overlooked in variant call format (VCF) files. TRGT-denovo improves the precision and specificity of de novo mutation (DNM) identification, reducing the number of de novo candidates by an order of magnitude compared to genotype-based approaches. In our experiments involving eight rare disease trios previously studied TRGT-denovo correctly reclassified all false positive DNM candidates as true negatives. Using an expanded repeat catalog, it identified new candidates, of which 95% (19/20) were experimentally validated, demonstrating its effectiveness in minimizing likely false positives while maintaining high sensitivity for true discoveries. Availability and implementationBuilt in Rust, TRGT-denovo is available as source code and a pre-compiled Linux binary along with a user guide at: https://github.com/PacificBiosciences/trgt-denovo.

bioinformatics↗

Resolving the unsolved: Comprehensive assessment of tandem repeats at scale

Tandem repeat (TR) variation is associated with gene expression changes and over 50 rare monogenic diseases. Recent advances in sequencing have enabled accurate, long reads that can characterize the full-length sequence and methylation profile of TRs. However, despite these advances in sequencing technology, computational methods to fully profile tandem repeats across the genome do not exist. To address this gap, we introduce tools for tandem repeat genotyping (TRGT), visualization and an accompanying TR database. TRGT accurately resolves the length and sequence composition of TR regions in the human genome. Assessing 937,122 TRs, TRGT showed a Mendelian concordance of 99.56%, allowing a single repeat unit difference. In six samples with known repeat expansions, TRGT detected all repeat expansions while also identifying methylation signals, mosaicism, and providing finer resolution of repeat length. Additionally, we release a database with allele sequences and methylation levels for 937,122 TRs across 100 genomes.

genomics↗