bioRxiv Science⌕ Search

Biology subjects

Porsborg, P. S.

Publications and source records attributed to Porsborg, P. S..

3 recordsLinked to original sources

Accurate calling of low-frequency somatic mutations by sample-specific modeling of error rates

Calling rare somatic variants from NGS data is more challenging than calling inherited variants, especially if the somatic variant is only present in a small fraction of the cells in the sequenced biopsy. In this case, having a good estimate of the error rate of a specific base in a particular read becomes essential. In paired-end sequencing, where some DNA fragments are shorter than twice the read length, the overlapping regions of the read pairs are an ideal resource for training models to discern context-dependent base error rates, as any discordant bases in the overlaps must be caused by a sequencing error or an alignment error. We have created a new tool named BBQ (an acronym for Better Base Quality) that uses overlapping reads to estimate the error rate conditional on the mutation type, sequence context, and base quality. We also estimate how much the error rate of concordant bases in overlapping reads is decreased compared to bases in non-overlapping reads. Results show that overlapping reads can remove sequencing errors induced by DNA damage and that the increased quality of overlapping reads differs between samples and mutation types, reflecting different damage patterns between samples. We use the error models to call rare somatic variants. Sequencing data from a testis biopsy and a cell-free DNA sample serve as a proof-of-concept for rare germ cell mutation calling and for detecting rare cancer mutations. We find that using the sample-specific error models of BBQ allows us to call rare somatic variants with fewer false positives than existing tools such as Mutect2 and Strelka2.

bioinformatics↗

Estimating gene conversion tract length and rate from PacBio HiFi data

Gene conversions are broadly defined as the transfer of genetic material from a donor to an acceptor sequence and can happen both in meiosis and mitosis. They are a subset of non-crossover events and, like crossover events, gene conversion can generate new combinations of alleles and counteract mutation load by reverting germline mutations through GC-biased gene conversion. Estimating gene conversion rate and the distribution of gene conversion tract lengths remains challenging. We present a new method for estimating tract length, rate and detection probability of non-crossover events directly in HiFi PacBio long read data. The method can be used to make inference from sequencing of gametes from a single individual. The method is unbiased even under low single nucleotide variant (SNV) densities and does not necessitate any demographic or evolutionary assumptions. We test the accuracy and robustness of our method using simulated datasets where we vary length of tracts, number of tracts, the genomic SNV density and levels of correlation between SNV density and NCO event position. Our simulations show that under low SNV densities, like those found in humans, only a minute fraction ([~]2%) of NCO events are expected to become visible as gene conversions by moving at least one SNV. We finally illustrate our method by applying it to PacBio sequencing data from human sperm.

genomics↗

Insights into gene conversion and crossing-over processes from long-read sequencing of human, chimpanzee and gorilla testes and sperm

Homologous recombination rearranges genetic information during meiosis to generate new combinations of variants. Recombination also causes new mutations, affects the GC content of the genome and reduces selective interference. Here, we use HiFi long-read sequencing to directly detect crossover and gene conversion events from switches between the two haplotypes along single HiFi-reads from testis tissue of humans, chimpanzees and gorillas as well as human sperm samples. Furthermore, based on DNA methylation calls, we classify the cellular origin of reads to either somatic or germline cells in the testis tissue. We identify 1692 crossovers and 1032 gene conversions in nine samples and investigate their chromosomal distribution. Crossovers are more telomeric and correlate better with recombination maps than gene conversions. We show a strong concordance between a human double-strand break map and the human samples, but not for the other species, supporting different PRDM9-programmed double-strand break loci. We estimate the average gene conversion tract lengths to be similar and very short in all three species (means 40-100 bp, fitted well by a geometric distribution) and that 95-98% of non-crossover events do not involve tracts intersecting with polymorphism and are therefore not detectable. Finally, we detect a GC bias in the gene conversion of both single and multiple SNVs and show that the GC-biased gene conversion affects SNVs flanking crossover events. This implies that gene conversion events associated with crossover events are much longer (estimated above 500 bp) than those associated with non-crossover events. Highly accurate long-read sequencing combined with the classification of reads to specific cell types provides a new, powerful way to make individual, detailed maps of gene conversion and crossovers for any species.

evolutionary biology↗