bioRxiv ScienceSearch

Biology subjects

Anthony T Papenfuss

Publications and source records attributed to Anthony T Papenfuss.

3 recordsLinked to original sources

Enrich2: a statistical framework for analyzing deep mutational scanning data

Measuring the functional consequences of protein variants can reveal how a protein works or help unlock the meaning of an individuals genome. Deep mutational scanning is a widely used method for multiplex measurement of the functional consequences of protein variants. A major limitation of this method has been the lack of a common analysis framework. We developed a statistical model for estimating variant scores that can be applied to many experimental designs. Our method generates an error estimate for each score that captures both sampling error and consistency between replicates. We apply our model to one novel and five published datasets comprising 243,732 variants and demonstrate its superiority, particularly for removing noisy variants, detecting variants of small effect, and conducting hypothesis testing. We implemented our model in easy-to-use software, Enrich2, that can empower researchers analyzing deep mutational scanning data.

Bioinformatics

De novo transcriptome assembly for the spiny mouse (Acomys cahirinus)

Background: Spiny mice of the genus Acomys are small desert-dwelling rodents that display physiological characteristics not typically found in rodents. Recent investigations have reported a menstrual cycle and scar free-wound healing in this species; characteristics that are exceedingly rare in mammals, and of considerable interest to the scientific community. These unique physiological traits, and the potential for spiny mice to accurately model human diseases, are driving increased use of this genus in biomedical research. However, little genetic information is currently available for Acomys, limiting the application of some modern investigative techniques. This project aimed to generate a reference transcriptome assembly for the common spiny mouse (Acomys cahirinus).\n\nResults: Illumina RNA sequencing of male and female spiny mice produced 451 million, 150bp paired-end reads from 15 organ types. An extensive survey of de novo transcriptome assembly approaches of high-quality reads using Trinity, SOAPdenovo-Trans, and Velvet/Oases at multiple kmer lengths was conducted with 49 single-kmer assemblies generated from this dataset, with and without in silico normalization and probabilistic error correction. Merging transcripts from 49 individual single-kmer assemblies into a single meta-assembly of non-redundant transcripts using the EvidentialGene tr2aacds pipeline produced the highest quality transcriptome assembly, comprised of 880,080 contigs, of which 189,925 transcripts were annotated using the SwissProt/Uniprot database.\n\nConclusions: This study provides the first detailed characterization of the spiny mouse transcriptome. It validates the application of the EvidentialGene tr2aacds pipeline to generate a high-quality reference transcriptome assembly in a mammalian species, and provides a valuable scientific resource for further investigation into the unique physiological characteristics inherent in the genus Acomys.

Bioinformatics

Improving the Power of Structural Variation Detection by Augmenting the Reference

The uses of the Genome Reference Consortiums human reference sequence can be roughly categorized into three related but distinct categories: as a representative species genome, as a coordinate system for identifying variants, and as an alignment reference for variation detection algorithms. However, the use of this reference sequence as simultaneously a representative species genome and as an alignment reference leads to unnecessary artifacts for structural variation detection algorithms and limits their accuracy. We show how decoupling these two references and developing a separate alignment reference can significantly improve the accuracy of structural variation detection, lead to improved genotyping of disease related genes, and decrease the cost of studying polymorphism in a population.

Bioinformatics