bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

ViR: a tool to account for intrasample variability in the detection of viral integrations

Lateral gene transfer (LT) from viruses to eukaryotic cells is a well-recognized phenomenon. Somatic integrations of viruses have been linked to persistent viral infection and genotoxic effects, including various types of cancer. As a consequence, several bioinformatic tools have been developed to identify viral sequences integrated into the human genome. Viral sequences that integrate into germline cells can be transmitted vertically, be maintained in host genomes and be co-opted for host functions. Endogenous viral elements (EVEs) have long been known, but the extent of their widespread occurrence has only been recently appreciated. Modern genomic sequencing analyses showed that eukaryotic genomes may harbor hundreds of EVEs, which derive not only from DNA viruses and retroviruses, but also from nonretroviral RNA viruses and are mostly enriched in repetitive regions of the genome. Despite being increasingly recognized as important players in different biological processes such as regulation of expression and immunity, the study of EVEs in non-model organisms has rarely gone beyond their characterization from annotated reference genomes because of the lack of computational methods suited to solve signals for EVEs in repetitive DNA. To fill this gap, we developed ViR, a pipeline which ameliorates the detection of integration sites by solving the dispersion of reads in genome assemblies that are rich of repetitive DNA. Using paired-end whole genome sequencing (WGS) data and a user-built database of viral genomes, ViR selects the best candidate couples of reads supporting an integration site by solving the dispersion of reads resulting from intrasample variability. We benchmarked ViR to work with sequencing data from both single and pooled DNA samples and show its applicability using WGS data of a non-model organism, the arboviral vector Aedes albopictus. Viral integrations predicted by ViR were molecularly validated supporting the accuracy of ViR results. Additionally, ViR can be readily adopted to detect any LT event providing ad hoc non-host sequences to interrogate.

bioinformatics↗

Transcriptogram analysis reveals relationship between viral titer and gene sets responses during Corona-virus infection.

To understand the difference between benign and severe outcomes after Coronavirus infection, we urgently need ways to clarify and quantify the time course of tissue and immune responses. Here we re-analyze 72-hour time-series microarrays generated in 2013 by Sims and collaborators for SARS-CoV-1 in vitro infection of a human lung epithelial cell line. Transcriptograms, a Bioinformatics tool to analyze genome-wide gene expression data, allow us to define an appropriate context-dependent threshold for mechanistic relevance of gene differential expression. Without knowing in advance which genes are relevant, classical analyses detect every gene with statistically-significant differential expression, leaving us with too many genes and hypotheses to be useful. Using a Transcriptogram-based top-down approach, we identified three major, differentially-expressed gene sets comprising 219 mainly immune-response-related genes. We identified timescales for alterations in mitochondrial activity, signaling and transcription regulation of the innate and adaptive immune systems and their relationship to viral titer. At the individual-gene level, EGR3 was significantly upregulated in infected cells. Similar activation in T-cells and fibroblasts in infected lung could explain the T-cell anergy and eventual fibrosis seen in SARS-CoV-1 infection. The methods can be applied to RNA data sets for SARS-CoV-2 to investigate the origin of differential responses in different tissue types, or due to immune or preexisting conditions or to compare cell culture, organoid culture, animal models, and human-derived samples.

bioinformatics↗

A graph-based approach identifies dynamic H-bond communication networks in spike protein S of SARS-CoV-2

Corona virus spike protein S is a large homo-trimeric protein embedded in the membrane of the virion particle. Protein S binds to angiotensin-converting-enzyme 2, ACE2, of the host cell, followed by proteolysis of the spike protein, drastic protein conformational change with exposure of the fusion peptide of the virus, and entry of the virion into the host cell. The structural elements that govern conformational plasticity of the spike protein are largely unknown. Here, we present a methodology that relies upon graph and centrality analyses, augmented by bioinformatics, to identify and characterize large H-bond clusters in protein structures. We apply this methodology to protein S ectodomain and find that, in the closed conformation, the three protomers of protein S bring the same contribution to an extensive central network of H-bonds, has a relatively large H-bond cluster at the receptor binding domain, and a cluster near a protease cleavage site. Markedly different H-bonding at these three clusters in open and pre-fusion conformations suggest dynamic H-bond clusters could facilitate structural plasticity and selection of a protein S protomer for binding to the host receptor, and proteolytic cleavage. From analyses of spike protein sequences we identify patches of histidine and carboxylate groups that could be involved in transient proton binding.

bioinformatics↗

Structure-aware Protein Solubility Prediction From Sequence Through Graph Convolutional Network And Predicted Contact Map

MotivationProtein solubility is significant in producing new soluble proteins that can reduce the cost of biocatalysts or therapeutic agents. Therefore, a computational model is highly desired to accurately predict protein solubility from the amino acid sequence. Many methods have been developed, but they are mostly based on the one-dimensional embedding of amino acids that is limited to catch spatially structural information. ResultsIn this study, we have developed a new structure-aware method to predict protein solubility by attentive graph convolutional network (GCN), where the protein topology attribute graph was constructed through predicted contact maps from the sequence. GraphSol was shown to substantially out-perform other sequence-based methods. The model was proven to be stable by consistent R2 of 0.48 in both the cross-validation and independent test of the eSOL dataset. To our best knowledge, this is the first study to utilize the GCN for sequence-based predictions. More importantly, this architecture could be extended to other protein prediction tasks. AvailabilityThe package is available at http://biomed.nscc-gz.cn Contactyangyd25@mail.sysu.edu.cn Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Ontology-guided segmentation and object identification for developmental mouse lung immunofluorescent images

BackgroundImmunofluorescent confocal microscopy uses labeled antibodies as probes against specific macromolecules to discriminate between multiple cell types. For images of the developmental mouse lung, these cells are themselves organized into densely packed higher-level anatomical structures. These types of images can be challenging to segment automatically for several reasons, including the relevance of biomedical context, dependence on the specific set of probes used, prohibitive cost of generating labeled training data, as well as the complexity and dense packing of anatomical structures in the image. The use of an application ontology surmounts these challenges by combining image data with its metadata to provide a meaningful biological context, and hence constraining and simplifying the process of segmentation and object identification. ResultsWe propose an innovative approach for the automated analysis of complex and densely packed anatomical structures from immunofluorescent images that utilizes an application ontology to provide a simplified context for image segmentation and object identification. We describe how the logical organization of biological facts in the form of an ontology can provide useful constraints that enhance automatic processing of complex images. We demonstrate the results of ontology-guided segmentation and object identification in mouse developmental lung images from the Bioinformatics REsource ATlas for the Healthy lung (BREATH) database of the Molecular Atlas of Lung Development (LungMAP1) program. ConclusionThe microscopy analysis pipeline library (micap) is available at https://github.com/duke-lungmap-team/microscopy-analysis-pipeline. Code to reproduce our analysis of LungMAP images is also available at https://github.com/duke-lungmap-team/lungmap-pipeline. Finally, the application ontology is available at https://github.com/duke-lungmap-team/lung_ontology and includes example SPARQL queries. ContactAnna Maria Masci email: annamaria.masci@duke.edu

bioinformatics↗

MetaFunPrimer: primer design for targeting genes observed in metagenomes

High throughput primer design is needed to simultaneously design primers for multiple genes of interest, such as a group of functional genes. We have developed MetaFunPrimer, a bioinformatic pipeline to design primer targets for genes of interests, with a prioritization based on ranking the presence of gene targets in references, such as metagenomes. MetaFunPrimer takes inputs of protein and nucleotide sequences for gene targets of interest accompanied by a set of reference metagenomes or genomes for determining genes of interest. Its output is a set of primers that may be used to amplify genes of interest. To demonstrate the usage and benefits of MetaFunPrimer, a total of 78 HT-qPCR primer pairs were designed to target observed ammonia monooxygenase subunit A (amoA) genes of ammonia-oxidizing bacteria (AOB) in 1,550 soil metagenomes. We demonstrate that these primers can significantly improve targeting of amoA-AOB genes in soil metagenomes compared to previously published primers. IMPORTANCEAmplification-based gene characterization allows for sensitive and specific quantification of functional genes. Often, there is a large diversity of genes represented for a function of interest, and multiple primers may be necessary to target associated genes. Current primer design tools are limited to designing primers for only a few genes of interest. MetaFunPrimer allows for high throughput primer design for functional genes of interest and also allows for ranking gene targets by their presence and abundance in environmental datasets. This tool enables high throughput qPCR approaches for characterizing functional genes.

bioinformatics↗

Renal carcinoma is associated with increased risk of coronavirus infections

The current pandemic COVID-19 has affected most severely to the people with old age, or with comorbidities such as hypertension, diabetes mellitus, chronic kidney disease, COPD, and cancers. Cancer patients are twice more likely to contract the disease because of the malignancy or treatment-related immunosuppression; hence identification of the vulnerable population among these patients is essential. It is speculated that along with ACE2, other auxiliary proteins (DPP4, ANPEP, ENPEP, TMPRSS2) might facilitate the entry of coronaviruses in the host cells. We took a bioinformatics approach to analyze the gene and protein expression data of these coronavirus receptors in human normal and cancer tissues of multiple organs. Here, we demonstrated an extensive RNA and protein expression profiling analysis of these receptors across solid tumors and normal tissues. We found that among all, renal tumor and normal tissues exhibited increased levels of ACE2, DPP4, ANPEP, and ENPEP. Our results revealed that TMPRSS2 may not be the co-receptor for coronavirus in renal carcinoma patients. The receptors’ expression levels were variable in different tumor stage, molecular and immune subtypes of renal carcinoma. In clear cell renal cell carcinomas, coronavirus receptors were associated with high immune infiltration, markers of immunosuppression, and T cell exhaustion. Our study indicates that CoV receptors may play an important role in modulating the immune infiltrate and hence cellular immunity in renal carcinoma. As our current knowledge of pathogenic mechanisms will improve, it may help us in designing focused therapeutic approaches.Competing Interest StatementThe authors have declared no competing interest.View Full Text

bioinformatics↗

The potential role of miR-21-3p in coronavirus-host interplay

ABSTRACTHost miRNAs are known as important regulators of virus replication and pathogenesis. They can interact with various viruses by several possible mechanisms including direct binding the viral RNA. Identification of human miRNAs involved in coronavirus-host interplay is becoming important due to the ongoing COVID-19 pandemic. In this work we performed computational prediction of high-confidence direct interactions between miRNAs and seven human coronavirus RNAs. In order to uncover the entire miRNA-virus interplay we further analyzed lungs miRNome of SARS-CoV infected mice using publicly available miRNA sequencing data. We found that miRNA miR-21-3p has the largest probability of binding the human coronavirus RNAs and being dramatically up-regulated in mouse lungs during infection induced by SARS-CoV. Further bioinformatic analysis of binding sites revealed high conservativity of miR-21-3p binding regions within RNAs of human coronaviruses and their strains.Competing Interest StatementThe authors have declared no competing interest.View Full Text

bioinformatics↗

ADMIXPIPE: Population analyses in ADMIXTURE for non-model organisms

Background: Research on the molecular ecology of non-model organisms, while previously constrained, has now been greatly facilitated by the advent of reduced-representation sequencing protocols. However, tools that allow these large datasets to be efficiently parsed are often lacking, or if indeed available, then limited by the necessity of a comparable reference genome as an adjunct. This, of course, can be difficult when working with non-model organisms. Fortunately, pipelines are currently available that avoid this prerequisite, thus allowing data to be a priori parsed. An oft-used molecular ecology program (i.e., STRUCTURE), for example, is facilitated by such pipelines, yet they are surprisingly absent for a second program that is similarly popular and computationally more efficient (i.e., ADMIXTURE). The two programs differ in that ADMIXTURE employs a maximum-likelihood framework whereas STRUCTURE uses a Bayesian approach, yet both produce similar results. Given these issues, there is an overriding (and recognized) need among researchers in molecular ecology for bioinformatic software that will not only condense output from replicated ADMIXTURE runs, but also infer from these data the optimal number of population clusters (K). Results: Here we provide such a program (i.e., ADMIXPIPE) that (a) filters SNPs to allow the delineation of population structure in ADMIXTURE, then (b) parses the output for summarization and graphical representation via CLUMPAK. Our benchmarks effectively demonstrate how efficient the pipeline is for processing large, non-model datasets generated via double digest restriction-site associated DNA sequencing (ddRAD). Outputs not only parallel those from STRUCTURE, but also visualize the variation among individual ADMIXTURE runs, so as to facilitate selection of the most appropriate K-value. Conclusions: ADMIXPIPE successfully integrates ADMIXTURE analysis with popular variant call format (VCF) filtering software to yield file types readily analyzed by CLUMPAK. Large population genomic datasets derived from non-model organisms are efficiently analyzed via the parallel-processing capabilities of ADMIXTURE. ADMIXPIPE is distributed under the GNU Public License and freely available for Mac OSX and Linux platforms at: https://github.com/stevemussmann/admixturePipeline.Competing Interest StatementThe authors have declared no competing interest.View Full Text

bioinformatics↗

PTM-Shepherd: analysis and summarization of post-translational and chemical modifications from open search results

Open searching has proven to be an effective strategy for identifying both known and unknown modifications in shotgun proteomics experiments. Rather than being limited to a small set of user-specified modifications, open searches identify peptides with any mass shift that may correspond to a single modification or a combination of several modifications. Here we present PTM-Shepherd, a bioinformatics tool that automates characterization of PTM profiles detected in open searches based on attributes such as amino acid localization, fragmentation spectra similarity, retention time shifts, and relative modification rates. PTM-Shepherd can also perform multi-experiment comparisons for studying changes in modification profiles, e.g. in data generated in different laboratories or under different conditions. We demonstrate how PTM-Shepherd improves the analysis of data from formalin-fixed paraffin-embedded samples, detects extreme underalkylation of cysteine in some datasets, discovers an artefactual modification introduced during peptide synthesis, and uncovers site-specific biases in sample preparation artifacts in a multi-center proteomics profiling study.

bioinformatics↗

DeepCDR: a hybrid graph convolutional network for predicting cancer drug response

Motivation Accurate prediction of cancer drug response (CDR) is challenging due to the uncertainty of drug efficacy and heterogeneity of cancer patients. Strong evidences have implicated the high dependence of CDR on tumor genomic and transcriptomic profiles of individual patients. Precise identification of CDR is crucial in both guiding anti-cancer drug design and understanding cancer biology.Results In this study, we present DeepCDR which integrates multi-omics profiles of cancer cells and explores intrinsic chemical structures of drugs for predicting cancer drug response. Specifically, DeepCDR is a hybrid graph convolutional network consisting of a uniform graph convolutional network (UGCN) and multiple subnetworks. Unlike prior studies modeling hand-crafted features of drugs, DeepCDR automatically learns the latent representation of topological structures among atoms and bonds of drugs. Extensive experiments showed that DeepCDR outperformed state-of-the-art methods in both classification and regression settings under various data settings. We also evaluated the contribution of different types of omics profiles for assessing drug response. Furthermore, we provided an exploratory strategy for identifying potential cancer-associated genes concerning specific cancer types. Our results highlighted the predictive power of DeepCDR and its potential translational value in guiding disease-specific drug design.Availability DeepCDR is freely available at https://github.com/kimmo1019/DeepCDRContact ruijiang{at}tsinghua.edu.cn; muzhou{at}sensebrain.siteSupplementary information Supplementary data are available at Bioinformatics online.Competing Interest StatementThe authors have declared no competing interest.View Full Text

bioinformatics↗

CRAFT: Compact genome Representation towards largescale Alignment-Free daTabase

MotivationRapid developments in sequencing technologies have boosted generating high volumes of sequence data. To archive and analyze those data, one primary step is sequence comparison. Alignment-free sequence comparison based on k-mer frequencies offers a computationally efficient solution, yet in practice, the k-mer frequency vectors for large k of practical interest lead to excessive memory and storage consumption. ResultsWe report CRAFT, a general genomic/metagenomic search engine to learn compact representations of sequences and perform fast comparison between DNA sequences. Specifically, given genome or high throughput sequencing (HTS) data as input, CRAFT maps the data into a much smaller embedding space and locates the best matching genome in the archived massive sequence repositories. With 102 - 104-fold reduction of storage space, CRAFT performs fast query for gigabytes of data within seconds or minutes, achieving comparable performance as six state-of-the-art alignment-free measures. AvailabilityCRAFT offers a user-friendly graphical user interface with one-click installation on Windows and Linux operating systems, freely available at https://github.com/jiaxingbai/CRAFT. Contactwangying@xmu.edu.cn; fsun@usc.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing

Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models taken from NLP. These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The LMs were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw protein LM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks. The first was a per-residue prediction of protein secondary structure (3-state accuracy Q3=81%-87%); the second were per-protein predictions of protein sub-cellular localization (ten-state accuracy: Q10=81%) and membrane vs. water-soluble (2-state accuracy Q2=91%). For the per-residue predictions the transfer of the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without using evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that protein LMs learned some of the grammar of the language of life. To facilitate future work, we released our models at https://github.com/agemagician/ProtTrans.

bioinformatics↗

Number of mismatches and length of longest match correlate with alignment score in swalign built-in function in MATLAB

Understanding how one sequence relates to another at the nucleotide or amino acid level allows the derivation of new knowledge regarding the provenance of particular sequence as well as the determination of consensus sequence motifs that informs biological conservation at the sequence level. To this end, local or multiple sequence alignments tools in bioinformatics have been developed to automatically profile two or more nucleotide or amino acid sequence in search of matches in stretches of nucleotides or amino acid sequence that yield an alignment. While alignment score is a common metric for assessing alignment quality, relative difference between alignment scores does not readily correlate with concrete measures such as number of mismatches and length of longest match in alignment. Thus, using swalign local sequence alignment function in MATLAB on 200 alignments between RNA-seq sequence read and reference Escherichia coli K-12 MG1655 genome sequence in the sense and antisense direction, this work sought to shed some light on how alignment score from swalign correlates with number of mismatches and length of longest match. Results revealed that number of mismatches negatively correlate with alignment score; thereby, validating theoretical predictions that larger number of mismatches would result in a poorer alignment and lower alignment score. However, dependence of alignment score on other factors such as length of longest match and gap penalty from opening an alignment gap prevents linear relationship to be obtained between number of mismatches and alignment score. On the other hand, length of longest match was found to positively correlate with alignment score as predicted from theoretical understanding. But, data obtained revealed that clusters of data points gather at two regions of the scatter plot involving short matches and low alignment score, as well as long matches and high alignment score. Such clustering and sparseness of data points between the two clusters preclude the elucidation of a linear quantitative relationship between length of longest match and alignment score. Overall, dependence of alignment score of swalign on number of mismatches and length of longest match in alignment match theoretical predictions; thereby, validating the utility of alignment score in indicating the qualitative quality of alignment. However, given that alignment score inherently depends on a multitude of factors, users could not easily discern the quantitative difference in mismatches and length of longest match from relative differences between two alignment scores. Such problems are unlikely to be resolved given the near impossibility of obtaining quantitative linear relationship correlating either number of mismatches or length of longest match with alignment score of a sequence alignment tool. HighlightsO_LINumber of mismatches in alignment negatively correlates with alignment score. C_LIO_LILength of longest match positively correlates with alignment score. C_LIO_LIQuantitative linear relationship could not be obtained for alignment score with either number of mismatches or length of longest match. C_LIO_LIResults validate that swalign tool in MATLAB could quantitatively detect differences in alignment quality and expressed it using alignment score. C_LIO_LIBut, relative alignment score of two alignments remains a nebulous concept with regards to differences in number of mismatches and length of longest match. C_LI

bioinformatics↗

Curation of over 10,000 transcriptomic studies to enable data reuse

Vast amounts of transcriptomic data reside in public repositories, but effective reuse remains challenging. Issues include unstructured dataset metadata, inconsistent data processing and quality control, and inconsistent probe-gene mappings across microarray technologies. Thus, extensive curation and data reprocessing is necessary prior to any reuse. The Gemma bioinformatics system was created to help address these issues. Gemma consists of a database of curated transcriptomic datasets, analytical software, a web interface, and web services. Here we present an update on Gemmas holdings, data processing and analysis pipelines, our curation guidelines, and software features. As of June 2020, Gemma contains 10,811 manually curated datasets (primarily human, mouse, and rat), over 395,000 samples and hundreds of curated transcriptomic platforms (both microarray and RNA-sequencing). Dataset topics were represented with 10,215 distinct terms from 12 ontologies, for a total of 54,316 topic annotations (mean topics/dataset = 5.2). While Gemma has broad coverage of conditions and tissues, it captures a large majority of available brain-related datasets, accounting for 34% of its holdings. Users can access the curated data and differential expression analyses through the Gemma website, RESTful service, and an R package. Database URL: https://gemma.msl.ubc.ca/home.html

bioinformatics↗

Computational and Molecular Dynamics Simulation ApproachTo Analyze the Impact of XPD Gene Mutation on Protein Stability and Function

XPD acts as a functional helicase and aids in unwinding double helix around damaged DNA, leading to efficient DNA repair. Mutations of XPD give rise to DNA-repair deficiency diseases and cancer proneness. In this study, cancer-causing missense mutation that could inactivate helicase function and hinder its binding with other complexes were analysed using bioinformatics approach. Rigorous computational methods were employed to understand the molecular pathogenic profile of mutation. The mutant model with the desired mutation was built with I-TASSER. GROMACS 5.0.1 was used to evaluate the effect of a mutation on protein stability and function. Of the 276 missense mutations, 64 were found to be disease-causing. Out of these 64, seven were of cancer-causing mutations. Among these, we evaluated K48R mutation in a computational simulated environment to determine its impact on protein stability and function since K48 position was ascertained to be highly conserved and substitution with arginine could impair the XPD activity. Molecular Dynamic Simulation and Essential Dynamics analysis showed that K48R mutation altered protein structural stability and produced conformational drift. Our predictions thus revealed that K48R mutation could impair the XPD helicase activity and affect its ability to repair the damaged DNA, thus augmenting the risk for cancer.

bioinformatics↗

A novel predicted ADP-ribosyltransferase family conserved in eukaryotic evolution

The presence of many completely uncharacterized proteins, even in well-studied organisms such as humans, seriously hampers full understanding of the functioning of the living cells. ADP-ribosylation is a common post-translational modification of proteins; also nucleic acids and small molecules can be modified by the covalent attachment of ADP-ribose. This modification, important in cellular signalling and infection processes, is usually executed by enzymes from the large superfamily of ADP-ribosyltransferases (ARTs) Here, using bioinformatics approaches, we identify a novel putative ADP-ribosyltransferase family, conserved in eukaryotic evolution, with a divergent active site. The hallmark of these proteins is the ART domain nestled between flanking leucine-rich repeat (LRR) domains. LRRs are involved in innate immune surveillance. The novel family appears as likely novel ADP-ribosylation "writers", previously unnoticed new players in cell signaling by this emerging post-translational modification. We propose that this family, including its human member LRRC9, may be involved in an ancient defense mechanism, with analogies to the innate immune system, and coupling pathogen detection to ADP-ribosyltransfer signalling.

bioinformatics↗

Assessment of current taxonomic assignment strategies for metabarcoding eukaryotes

The effective use of metabarcoding in biodiversity science has brought important analytical challenges due to the need to generate accurate taxonomic assignments. The assignment of sequences to a generic or species level is critical for biodiversity surveys and biomonitoring, but it is particularly challenging. Researchers must select the approach that best recovers information on species composition. This study evaluates the performance and accuracy of seven methods in recovering the species composition of mock communities which vary in species number and specimen abundance, while holding upstream molecular and bioinformatic variables constant. It also evaluates the impact of parameter optimization on the quality of the predictions. Despite the general belief that BLAST top hit underperforms newer methods, our results indicate that it competes well with more complex approaches if optimized for the mock community under study. For example, the two machine learning methods that were benchmarked proved more sensitive to the reference database heterogeneity and completeness than methods based on sequence similarity. The accuracy of assignments was impacted by both species and specimen counts which will influence the selection of appropriate software. We urge the usage of realistic mock communities to allow optimization of parameters, regardless of the taxonomic assignment method used.

bioinformatics↗