bioRxiv ScienceSearch

Biology subjects

Hoff, K. J.

Publications and source records attributed to Hoff, K. J..

2 recordsLinked to original sources

VARUS: Sampling Complementary RNA Reads from the Sequence Read Archive

Vast amounts of next generation sequencing RNA data has been deposited in archives, accompanying very diverse original studies. The data is readily available also for other purposes such as genome annotation or transcriptome assembly. However, selecting a subset of available experiments, sequencing runs and reads for this purpose is a nontrivial task and complicated by the inhomogeneity of the data.\n\nThis article presents the software VARUS that selects, downloads and aligns reads from NCBIs Sequence Read Archive, given only the species binomial name and genome. VARUS automatically chooses runs from among all archived runs to randomly select subsets of reads. The objective of its online algorithm is to cover a large number of transcripts adequately when network bandwidth and computing resources are limited. For most tested species VARUS achieved both a higher sensitivity and specificity with a lower number of downloaded reads than when runs were manually selected. At the example of twelve eukaryotic genomes, we show that RNA-Seq that was sampled with VARUS is well-suited for fully-automatic genome annotation with BRAKER.\n\nWith VARUS, genome annotation can be automatized to the extent that not even the selection and quality control of RNA-Seq has to be done manually. This introduces the possibility to have fully automatized genome annotation loops over potentially many species without incurring a loss of accuracy over a manually supervised annotation process.

bioinformatics

MakeHub: Fully automated generation of UCSC Genome Browser Assembly Hubs

Novel genomes are today often annotated by small consortia or individuals whose background is not from bioinformatics. This audience requires tools that are easy to use. This need had been addressed by several genome annotation tools and pipelines. Visualizing resulting annotation is a crucial step of quality control. The UCSC Genome Browser is a powerful and popular genome visualization tool. Assembly Hubs allow browsing genomes that are hosted locally via already available UCSC Genome Browser servers. The steps for creating custom Assembly Hubs are well documented and the required tools are publicly available. However, the number of steps for creating a novel Assembly Hub is large. In some cases the format of input files needs to be adapted which is a difficult task for scientists without programming background. Here, we describe the novel command line tool MakeHub that generates Assembly Hubs for the UCSC Genome Browser in a fully automated fashion. The pipeline also allows extending previously created Hubs by additional tracks. MakeHub is freely available for download from https://github.com/Gaius-Augustus/MakeHub. Contactkatharina.hoff@uni-greifswald.de

bioinformatics