bioRxiv ScienceSearch

Biology subjects

Steinegger, M.

Publications and source records attributed to Steinegger, M..

3 recordsLinked to original sources

MMseqs2 desktop and local web server app for fast, interactive sequence searches

The MMseqs2 desktop and web server app facilitates interactive sequence searches through custom protein sequence and profile databases on personal workstations. By eliminating MMseqs2s runtime overhead, we reduced response times to a few seconds at sensitivities close to BLAST.\n\nAvailability and implementationThe app is easy to install for non-experts. Source code, prebuilt desktop app packages for Windows, macOS and Linux, Docker images for the web server application, and a demo web server are available at https://search.mmseqs.com.\n\nContactmartin.steinegger@mpibpc.mpg.de or soeding@mpibpc.mpg.de

bioinformatics

Protein-level assembly increases protein sequence recovery from metagenomic samples manyfold

The open-source de-novo Protein-level assembler Plass (https://plass.mmseqs.org) assembles six-frame-translated sequencing reads into protein sequences. It recovers 2 to 10 times more protein sequences from complex metagenomes and can assemble huge datasets. We assembled two redundancy-filtered reference protein catalogs, 2 billion sequences from 640 soil samples (SRC) and 292 million sequences from 775 marine eukaryotic metatranscriptomes (MERC), the largest free collections of protein sequences.

bioinformatics

Linclust: clustering protein sequences in linear time

Metagenomic datasets contain billions of protein sequences that could greatly enhance large-scale functional annotation and structure prediction. Utilizing this enormous resource would require reducing its redundancy by similarity clustering. However, clustering hundreds of million of sequences is impractical using current algorithms because their runtimes scale as the input set size N times the number of clusters K, which is typically of similar order as N, resulting in runtimes that increase almost quadratically with N. We developed Linclust, the first clustering algorithm whose runtime scales as N, independent of K. It can also cluster datasets several times larger than the available main memory. We cluster 1.6 billion metagenomic sequence fragments in 10 hours on a single server to 50% sequence identity, > 1000 times faster than has been possible before. Linclust will help to unlock the great wealth contained in metagenomic and genomic sequence databases. (Open-source software and Metaclust database: https://mmseqs.org/).

bioinformatics