bioRxiv ScienceSearch

Biology subjects

Starrett, G. J.

Publications and source records attributed to Starrett, G. J..

3 recordsLinked to original sources

Clinical and molecular characterization of virus-positive and virus-negative Merkel cell carcinoma

Merkel cell carcinoma (MCC) is a highly aggressive neuroendocrine carcinoma of the skin mediated by the integration of Merkel cell polyomavirus (MCPyV) and expression of viral T antigens or by ultraviolet induced damage to the tumor genome from excessive sunlight exposure. An increasing number of deep sequencing studies of MCC have identified significant differences between the number and types of point mutations, copy number alterations, and structural variants between virus-positive and virus-negative tumors. In this study, we assembled a cohort of 71 MCC patients and performed deep sequencing with OncoPanel, a next-generation sequencing assay targeting over 400 cancer-associated genes. To improve the accuracy and sensitivity for virus detection compared to traditional PCR and IHC methods, we developed a hybrid capture baitset against the entire MCPyV genome. The viral baitset identified integration junctions in the tumor genome and generated assemblies that strongly support a model of a hybrid, virus-host, circular DNA intermediate during integration that promotes focal amplification of host DNA. Using the clear delineation between virus-positive and virus-negative tumors from this method, we identified recurrent somatic alterations common across MCC and alterations specific to each class of tumor, associated with differences in overall survival. Comparing the molecular and clinical data from these patients revealed a surprising association of immunosuppression with virus-negative MCC and significantly shortened overall survival. These results demonstrate the value of high-confidence virus detection for identifying clinically important features in MCC that impact patient outcome.

cancer biology

Mash Screen: High-throughput sequence containment estimation for genome discovery

The MinHash algorithm has proven effective for rapidly estimating the resemblance of two genomes or metagenomes. However, this method cannot reliably estimate the containment of a genome within a metagenome. Here we describe an online algorithm capable of measuring the containment of genomes and proteomes within either assembled or unassembled sequencing read sets. We describe several use cases, including contamination screening and retrospective analysis of metagenomes for novel genome discovery. Using this tool, we provide containment estimates for every NCBI RefSeq genome within every SRA metagenome, and demonstrate the identification of a novel polyomavirus species from a public metagenome.

bioinformatics

Discovery of several thousand highly diverse circular DNA viruses

Although it is suspected that there are millions of distinct viral species, fewer than 9,000 are catalogued in GenBanks RefSeq database. We selectively enriched for and amplified the genomes of circular DNA viruses in over 70 animal samples, ranging from cultured soil nematodes to human tissue specimens. A bioinformatics pipeline, Cenote-Taker, was developed to automatically annotate over 2,500 circular genomes in a GenBank-compliant format. The new genomes belong to dozens of established and emerging viral families. Some appear to be the result of previously undescribed recombination events between ssDNA viruses and ssRNA viruses. In addition, hundreds of circular DNA elements that do not encode any discernable similarities to previously characterized sequences were identified. To characterize these "dark matter" sequences, we used an artificial neural network to identify candidate viral capsid proteins, several of which formed virus-like particles when expressed in culture. These data further the understanding of viral sequence diversity and allow for high throughput documentation of the virosphere.

microbiology