bioRxiv Science⌕ Search

Biology subjects

Oroperv, C.

Publications and source records attributed to Oroperv, C..

2 recordsLinked to original sources

A reference-free strategy for circulating tumor DNA detection from whole-genome sequencing data

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed reference-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmers tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

bioinformatics↗

Accurate calling of low-frequency somatic mutations by sample-specific modeling of error rates

Calling rare somatic variants from NGS data is more challenging than calling inherited variants, especially if the somatic variant is only present in a small fraction of the cells in the sequenced biopsy. In this case, having a good estimate of the error rate of a specific base in a particular read becomes essential. In paired-end sequencing, where some DNA fragments are shorter than twice the read length, the overlapping regions of the read pairs are an ideal resource for training models to discern context-dependent base error rates, as any discordant bases in the overlaps must be caused by a sequencing error or an alignment error. We have created a new tool named BBQ (an acronym for Better Base Quality) that uses overlapping reads to estimate the error rate conditional on the mutation type, sequence context, and base quality. We also estimate how much the error rate of concordant bases in overlapping reads is decreased compared to bases in non-overlapping reads. Results show that overlapping reads can remove sequencing errors induced by DNA damage and that the increased quality of overlapping reads differs between samples and mutation types, reflecting different damage patterns between samples. We use the error models to call rare somatic variants. Sequencing data from a testis biopsy and a cell-free DNA sample serve as a proof-of-concept for rare germ cell mutation calling and for detecting rare cancer mutations. We find that using the sample-specific error models of BBQ allows us to call rare somatic variants with fewer false positives than existing tools such as Mutect2 and Strelka2.

bioinformatics↗