bioRxiv Science⌕ Search

Biology subjects

Andersen, C. L.

Publications and source records attributed to Andersen, C. L..

4 recordsLinked to original sources

Error-corrected deep targeted sequencing of circulating cell-free DNA from colorectal cancer patients for sensitive detection of circulating tumor DNA

IntroductionCirculating tumor DNA (ctDNA) is a promising biomarker, reflecting the presence of tumor cells. Sequencing-based detection of ctDNA at low tumor fractions is challenging due to the crude error rate of sequencing. To mitigate this challenge, we developed UMIseq, a fixed-panel deep-targeted sequencing approach, which is universally applicable to all colorectal cancer (CRC) patients. MethodsUMIseq features UMI-mediated error correction, exclusion of mutations related to clonal hematopoiesis, a panel of normals for error modeling, and signal integration from single-nucleotide variations, insertions, deletions, and phased mutations. UMIseq was trained and independently validated on pre-operative (pre-OP) plasma from CRC patients (n=364) and healthy individuals (n=61). ResultsUMIseq displayed an area under the curve surpassing 0.95 for allele frequencies (AF) down to 0.05%. In the training cohort, the pre-OP detection rate reached 80% at 95% specificity, while in the validation cohort, it was 70%. UMIseq enabled the detection of AFs down to 0.004%. To assess the potential for detection of residual disease, 26 post-operative plasma samples from stage III CRC patients were analyzed. Detection of ctDNA was associated with recurrence (p =0.08). ConclusionUMIseq demonstrated robust performance with high sensitivity and specificity, enabling the detection of ctDNA at low allele frequencies.

cancer biology↗

DREAMS: Deep Read-level Error Model for Sequencing data applied to low-frequency variant calling and circulating tumor DNA detection

Circulating tumor DNA detection using Next-Generation Sequencing (NGS) data of plasma DNA is promising for cancer identification and characterization. However, the tumor signal in the blood is often low and difficult to distinguish from errors. We present DREAMS (Deep Read-level Modelling of Sequencing-errors) for estimating error rates of individual read positions. Using DREAMS, we developed statistical methods for variant calling (DREAMS-vc) and cancer detection (DREAMS-cc). For evaluation, we generated deep targeted NGS data of matching tumor and plasma DNA from 85 colorectal cancer patients. The DREAMS approach performed better than state-of-the-art methods for variant calling and cancer detection.

bioinformatics↗

Machine learning guided signal enrichment for ultrasensitive plasma tumor burden monitoring

In solid tumor oncology, circulating tumor DNA (ctDNA) is poised to transform care through accurate assessment of minimal residual disease (MRD) and therapeutic response monitoring. To overcome the sparsity of ctDNA fragments in low tumor fraction (TF) settings and increase MRD sensitivity, we previously leveraged genome-wide mutational integration through plasma whole genome sequencing (WGS). We now introduce MRD-EDGE, a composite machine learning-guided WGS ctDNA single nucleotide variant (SNV) and copy number variant (CNV) detection platform designed to increase signal enrichment. MRD-EDGE uses deep learning and a ctDNA-specific feature space to increase SNV signal to noise enrichment in WGS by 300X compared to our previous noise suppression platform MRDetect. MRD-EDGE also reduces the degree of aneuploidy needed for ultrasensitive CNV detection through WGS from 1Gb to 200Mb, thereby expanding its applicability to a wider range of solid tumors. We harness the improved performance to track changes in tumor burden in response to neoadjuvant immunotherapy in non-small cell lung cancer and demonstrate ctDNA shedding in precancerous colorectal adenomas. Finally, the radical signal to noise enrichment in MRD-EDGE enables de novo mutation calling in melanoma without matched tumor, yielding clinically informative TF monitoring for patients on immune checkpoint inhibition.

genomics↗

Discovering fragment length signatures of circulating tumor DNA using Non-negative Matrix Factorization

Sequencing of cell-free DNA (cfDNA) is currently being used to detect cancer by searching both for mutational and non-mutational alterations. Recent work has shown that the length distribution of cfDNA fragments from a cancer patient can inform tumor load and type. Here, we propose non-negative matrix factorization (NMF) of fragment length distributions as a novel method to study fragment lengths in cfDNA samples. Using shallow whole-genome sequencing (sWGS) of cfDNA from a cohort of patients with metastatic castration-resistant prostate cancer (mCRPC), we demonstrate how NMF accurately infers the true tumor fragment length distribution as an NMF component - and that the sample weights of this component correlate with ctDNA levels (r=0.75). We further demonstrate how using several NMF components enables accurate cancer detection on data from various early stage cancers (AUC=0.96). Finally, we show that NMF, when applied across genomic regions, can be used to discover fragment length signatures associated with open chromatin.

bioinformatics↗