bioRxiv ScienceSearch

Biology subjects

Kozlov, A. M.

Publications and source records attributed to Kozlov, A. M..

2 recordsLinked to original sources

ParGenes: a tool for massively parallel model selection and phylogenetic tree inference on thousands of genes.

MotivationCoalescent- and reconciliation-based methods are now widely used to infer species phylogenies from genomic data. They typically use per-gene phylogenies as input, which requires conducting multiple individual tree inferences on a large set of multiple sequence alignments (MSAs). At present, no easy-to-use parallel tool for this task exists. Ad hoc scripts for this purpose do not only induce additional implementation overhead, but can also lead to poor resource utilization and long times-to-solution. We present ParGenes, a tool for simultaneously determining the best-fit model and inferring maximum likelihood (ML) phylogenies on thousands of independent MSAs using supercomputers.\n\nResultsParGenes executes common phylogenetic pipeline steps such as model-testing, ML inference(s), bootstrapping, and computation of branch support values via a single parallel program invocation. We evaluated ParGenes by inferring > 20, 000 phylogenetic gene trees with bootstrap support values from Ensembl Compara and VectorBase alignments in 28 hours on a cluster with 1024 nodes.\n\nAvailabilityGNU GPL at https://github.com/BenoitMorel/ParGenes.\n\nContactBenoit.Morel@h-its.org\n\nSupplementary informationSupplementary material is available at Bioinformatics online.

bioinformatics

EPA-ng: Massively Parallel Evolutionary Placement of Genetic Sequences

Next Generation Sequencing (NGS) technologies have led to a ubiquity of molecular sequence data. This data avalanche is particularly challenging in metagenetics, which focuses on taxonomic identification of sequences obtained from diverse microbial environments. To achieve this, phylogenetic placement methods determine how these sequences fit into an evolutionary context. Previous implementations of phylogenetic placement algorithms, such as the Evolutionary Placement Algorithm (EPA) included in RAxML, or O_SCPLOWPPLACERC_SCPLOW, are being increasingly used for this purpose. However, due to the steady progress in NGS technologies, the current implementations face substantial scalability limitations. Here we present EPA-O_SCPLOWNGC_SCPLOW, a complete reimplementation of the EPA that is substantially faster, offers a distributed memory parallelization, and integrates concepts from both, RAxML-EPA, and O_SCPLOWPPLACERC_SCPLOW. EPA-O_SCPLOWNGC_SCPLOW can be executed on standard shared memory, as well as on distributed memory systems (e.g., computing clusters). To demonstrate the scalability of EPA-O_SCPLOWNGC_SCPLOW we placed 1 billion metagenetic reads from the Tara Oceans Project onto a reference tree with 3,748 taxa in just under 7 hours, using 2,048 cores. Our performance assessment shows that EPA-O_SCPLOWNGC_SCPLOW outperforms RAxML-EPA and O_SCPLOWPPLACERC_SCPLOW by up to a factor of 30 in sequential execution mode, while attaining comparable parallel efficiency on shared memory systems. We further show that the distributed memory parallelization of EPA-O_SCPLOWNGC_SCPLOW scales well up to 3,520 cores. EPA-O_SCPLOWNGC_SCPLOW is available under the AGPLv3 license: https://github.com/Pbdas/epa-ng

bioinformatics