bioRxiv Science⌕ Search

Biology subjects

Budowle, B.

Publications and source records attributed to Budowle, B..

4 recordsLinked to original sources

Precise and ultrafast tandem repeat variant detection in massively parallel sequencing reads

Calling tandem repeat (TR) variants from DNA sequences is of both theoretical and practical significance. A large number of software tools have been developed for detecting TRs. However, little study has been done to detect TR alleles from long-read sequences, and the effectiveness of detecting TR alleles from whole genome sequence (WGS) data still needs to be improved. Herein, a novel algorithm is described to retrieve TR regions from sequence alignment, and a software program, TRcaller, has been developed to call TR alleles from both short- and long-read sequences, both whole genome and targeted sequences generated from multiple sequencing platforms. The results showed that TRcaller could provide substantially higher accuracy in detecting TR alleles with magnitudes faster than the mainstream software tools. TRcaller is able to facilitate scalable, accurate, and ultrafast TR allele calling from large-scale sequence datasets in various applications, such as DNA forensics, medical research, disease diagnosis, evolution, and breeding programs. AvailabilityTRcaller is available at www.trcaller.com.

bioinformatics↗

USAT: a Bioinformatic Toolkit to Facilitate Interpretation and Comparative Visualization of Tandem Repeat Sequences

Tandem repeats (TR), which are highly variable genomic variants, are widely used in individual identification, disease diagnostics and evolutionary studies. The recent advances of sequencing technologies and bioinformatic tools facilitate calling TR haplotypes. Both length-based and sequence-based TR alleles are used in different applications. However, sequence-based TR alleles could provide the highest precision to characterize TR haplotypes. Analysis of the differences between or among TR haplotypes, especially at the single nucleotide level, is the focus of TR haplotype characterization. In this study, we developed a Universal STR Allele Toolkit (USAT) for TR haplotype analysis, which includes allele size conversion, sequence comparison of haplotypes, figure plotting and comparison for allele distribution, and interactive visualization. An example application of USAT for analysis of the CODIS core STR loci with benchmarking human individuals demonstrated the capabilities of USAT. USAT has a user-friendly graphic interface and runs in all major computing operating systems at a fast speed with parallel computing enabled. In summary, USAT is able to facilitate the interpretation, visualization, and comparisons of TRs.

bioinformatics↗

vcferr: Development, Validation, and Application of a SNP Genotyping Error Simulation Framework

MotivationGenotyping error can impact downstream SNP-based analyses. Simulating various modes and and levels of error can help investigators better understand potential biases caused by miscalled genotypes. ResultsWe have developed and validated vcferr, a tool to probabilistically simulate genotyping error and missigness in VCF files. We demonstrate how vcferr could be used to address a research question by introducing varying levels of error of different type into a sample in a simulated pedigree, and assessed how kinship analysis degrades as a function of kind and type of error. Software Availabilityvcferr is available for installation via PyPi (https://pypi.org/project/vcferr/) or conda (https://anaconda.org/bioconda/vcferr). The software is released under the MIT license with source code available on GitHub (https://github.com/signaturescience/vcferr).

bioinformatics↗

skater: An R package for SNP-based Kinship Analysis, Testing, and Evaluation

MotivationSNP-based kinship analysis with genome-wide relationship estimation and IBD segment analysis methods produces results that often require further downstream processing and manipulation. A dedicated software package that consistently and intuitively implements this analysis functionality is needed. ResultsHere we present the skater R package for SNP-based kinship analysis, testing, and evaluation with R. The skater package contains a suite of well-documented tools for importing, parsing, and analyzing pedigree data, performing relationship degree inference, benchmarking relationship degree classification, and summarizing IBD segment data. AvailabilityThe skater package is implemented as an R package and is released under the MIT license at https://github.com/signaturescience/skater. Documentation is available at https://signaturescience.github.io/skater.

bioinformatics↗