bioRxiv Science⌕ Search

Biology subjects

Costello, K.

Publications and source records attributed to Costello, K..

2 recordsLinked to original sources

The structural diversity of telomeres and centromeres across mouse subspecies revealed by complete assemblies

It is over twenty years since the publication of the C57BL/6J mouse reference genome, which has been a key catalyst for understanding mammalian disease biology. However, the mouse reference genome still lacks telomeres and centromeres, contains 281 chromosomal sequence gaps, and only partially represents many biomedically relevant loci. We present the first T2T mouse genomes for two key inbred strains, C57BL/6J and CAST/EiJ. These T2T genomes reveal significant variability in telomere and centromere sizes and structural organisation. We add an additional 213 Mbp of novel sequence to the reference genome containing 517 protein-coding genes. We examined two important but incomplete loci in the mouse genome - the pseudoautosomal region (PAR) on the sex chromosomes and KRAB zinc finger proteins (KZFPs) loci. We identified distant locations of the PAR boundary, different copy number and sizes of segmental duplications, and a multitude of amino acid substitution mutations in PAR genes.

genomics↗

Using Machine Learning to Facilitate Classification of Somatic Variants from Next-Generation Sequencing

BackgroundMolecular profiling has become essential for tumor risk stratification and treatment selection. However, cancer genome complexity and technical artifacts make identification of real variants a challenge. Currently, clinical laboratories rely on manual screening, which is costly, subjective, and not scalable. Here we present a machine learning-based method to distinguish artifacts from bona fide Single Nucleotide Variants (SNVs) detected by NGS from tumor specimens.\n\nMethodsA cohort of 11,278 SNVs identified through clinical sequencing of tumor specimens were collected and divided into training, validation, and test sets. Each SNV was manually inspected and labeled as either real or artifact as part of clinical laboratory workflow. A three-class (real, artifact and uncertain) model was developed on the training set, fine-tuned using the validation set, and then evaluated on the test set. Prediction intervals reflecting the certainty of the classifications were derived during the process to label \"uncertain\" variants.\n\nResultsThe optimized classifier demonstrated 100% specificity and 97% sensitivity over 5,587 SNVs of the test set. 1,252 out of 1,341 true positive variants were identified as real, 4,143 out of 4,246 false positive calls were deemed artifacts, while only 192(3.4%) SNVs were labeled as \"uncertain\" with zero misclassification between the true positives and artifacts in the test set.\n\nConclusionsWe presented a computational classifier to identify variant artifacts detected from tumor sequencing. Overall, 96.6% of the SNVs received a definitive label and thus were exempt from manual review. This framework could improve quality and efficiency of variant review process in clinical labs.

bioinformatics↗