bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,621 records · Page 90Linked to original sources

Assessment of current taxonomic assignment strategies for metabarcoding eukaryotes

The effective use of metabarcoding in biodiversity science has brought important analytical challenges due to the need to generate accurate taxonomic assignments. The assignment of sequences to a generic or species level is critical for biodiversity surveys and biomonitoring, but it is particularly challenging. Researchers must select the approach that best recovers information on species composition. This study evaluates the performance and accuracy of seven methods in recovering the species composition of mock communities which vary in species number and specimen abundance, while holding upstream molecular and bioinformatic variables constant. It also evaluates the impact of parameter optimization on the quality of the predictions. Despite the general belief that BLAST top hit underperforms newer methods, our results indicate that it competes well with more complex approaches if optimized for the mock community under study. For example, the two machine learning methods that were benchmarked proved more sensitive to the reference database heterogeneity and completeness than methods based on sequence similarity. The accuracy of assignments was impacted by both species and specimen counts which will influence the selection of appropriate software. We urge the usage of realistic mock communities to allow optimization of parameters, regardless of the taxonomic assignment method used.

bioinformatics↗

DeepG4 : A deep learning approach to predict active G-quadruplexes

DNA is a complex molecule carrying the instructions an organism needs to develop, live and reproduce. In 1953, Watson and Crick discovered that DNA is composed of two chains forming a double-helix. Later on, other structures of DNA were discovered and shown to play important roles in the cell, in particular G-quadruplex (G4). Following genome sequencing, several bioinformatic algorithms were developed to map G4s in vitro based on a canonical sequence motif, G-richness and G-skewness or alternatively sequence features including k-mers, and more recently machine/deep learning. Here, we propose a novel convolutional neural network (DeepG4) to map active G4s (forming both in vitro and in vivo). DeepG4 is very accurate to predict active G4s, while most state-of-the-art algorithms fail. Moreover, DeepG4 identifies key DNA motifs that are predictive of G4 activity. We found that active G4 motifs do not follow a very flexible sequence pattern as current algorithms seek for. Instead, active G4s are determined by numerous specific motifs. Moreover, among those motifs, we identified known transcription factors (TFs) which could play important roles in G4 activity by contributing either directly to G4 structures themselves or indirectly by participating in G4 formation in the vicinity. Moreover, we showed that specific TFs might explain G4 activity depending on cell type. Lastly, variant analysis suggests that SNPs altering predicted G4 activity could affect transcription and chromatin, e.g. gene expression, H3K4me3 mark and DNA methylation. Thus, DeepG4 paves the way for future studies assessing the impact of known disease-associated variants on DNA secondary structure by providing a mechanistic interpretation of SNP impact on transcription and chromatin. Availability: https://github.com/morphos30/DeepG4. Author summaryDNA is a molecule carrying genetic information and found in all living cells. In 1953, Watson and Crick found that DNA has a double helix structure. However, other DNA structures were later identified, and most notably, G-quadruplex (G4). In 2000, the Human Genome Project revealed the widespread presence of G4s in the genome using algorithms. To date, all G4 mapping algorithms were developed to map G4s on naked DNA, without knowing if they could be formed in the cell. Here, we designed a novel artificial intelligence algorithm that could map G4s active in the cell from the DNA sequence. We showed its better accuracy compared to existing algorithms. Moreover, we identified key transcriptional factor motifs that could explain G4 activity depending on cell type. Lastly, we demonstrated the existence of mutations that could alter G4 activity and therefore impact molecular processes, such as transcription, in the cell. Such results could provide a novel mechanistic interpretation of known disease-associated mutations.

bioinformatics↗

Accurate and sensitive detection of microbial eukaryotes from metagenomic shotgun sequencing data

BackgroundMicrobial eukaryotes are found alongside bacteria and archaea in natural microbial systems, including host-associated microbiomes. While microbial eukaryotes are critical to these communities, they are challenging to study with shotgun sequencing techniques and are therefore often excluded. ResultsHere we present EukDetect, a bioinformatics method to identify eukaryotes in shotgun metagenomic sequencing data. Our approach uses a database of 521,824 universal marker genes from 241 conserved gene families, which we curated from 3,713 fungal, protist, non-vertebrate metazoan, and non-streptophyte archaeplastid genomes and transcriptomes. EukDetect has a broad taxonomic coverage of microbial eukaryotes, performs well on low-abundance and closely related species, and is resilient against bacterial contamination in eukaryotic genomes. Using EukDetect, we describe the spatial distribution of eukaryotes along the human gastrointestinal tract, showing that fungi and protists are present in the lumen and mucosa throughout the large intestine. We discover that there is a succession of eukaryotes that colonize the human gut during the first years of life, mirroring patterns of developmental succession observed in gut bacteria. By comparing DNA and RNA sequencing of paired samples from human stool, we find that many eukaryotes continue active transcription after passage through the gut, though some do not, suggesting they are dormant or nonviable. We analyze metagenomic data from the Baltic Sea and find that eukaryotes differ across locations and salinity gradients. Finally, we observe eukaryotes in Arabidopsis leaf samples, many of which are not identifiable from public protein databases. ConclusionsEukDetect provides an automated and reliable way to characterize eukaryotes in shotgun sequencing datasets from diverse microbiomes. We demonstrate that it enables discoveries that would be missed or clouded by false positives with standard shotgun sequence analysis. EukDetect will greatly advance our understanding of how microbial eukaryotes contribute to microbiomes.

bioinformatics↗

HTSeqQC: A Flexible and One-Step Quality Control Software for High-throughput Sequence Data Analysis

Use of high-throughput sequencing (HTS) has become indispensable in life science research. Raw HTS data contains several sequencing artifacts, and as a first step it is imperative to remove the artifacts for reliable downstream bioinformatics analysis. Although there are multiple stand-alone tools available that can perform the various quality control steps separately, availability of an integrated tool that can allow one-step, automated quality control analysis of HTS datasets will significantly enhance handling large number of samples parallelly. Here, we developed HTSQualC, a stand-alone, flexible, and easy-to-use software for one-step quality control analysis of raw HTS data. HTSQualC can evaluate HTS data quality and perform filtering and trimming analysis in a single run. We evaluated the performance of HTSQualC for conducting batch analysis of HTS datasets with 322 samples with an average [~]1M (paired end) sequence reads per sample. HTSQualC accomplished the QC analysis in [~]3 hours in distributed mode and [~]31 hours in shared mode, thus underscoring its utility and robust performance. In addition to command-line execution, we integrated HTSQualC into the free, open-source, CyVerse cyberinfrastructure resource as a GUI interface, for wider access to experimental biologists who have limited computational resources and/or programming abilities.

bioinformatics↗

Evaluating the Number of Different Genomes in a Metagenome by Means of the Compositional Spectra Approach

Determination of metagenome composition is still one of the most interesting problems of bioinformatics. It involves a wide range of mathematical methods, from probabilistic models of combinatorics to cluster analysis and pattern recognition techniques. The successful advance of rapid sequencing methods and fast and precise metagenome analysis will increase the diagnostic value of healthy or pathological human metagenomes. The article presents the theoretical foundations of the algorithm for calculating the number of different genomes in the medium under study. The approach is based on analysis of the compositional spectra of subsequently sequenced samples of the medium. Its essential feature is using random fluctuations in the bacteria number in different samples of the same metagenome. The possibility of effective implementation of the algorithm in the presence of data errors is also discussed. In the work, the algorithm of a metagenome evaluation is described, including the estimation of the genome number and the identification of the genomes with known compositional spectra. It should be emphasized that evaluating the genome number in a metagenome can be always helpful, regardless of the metagenome separation techniques, such as clustering the sequencing results or marker analysis.

bioinformatics↗

Integrating structure-based machine learning and co-evolution to investigate specificity in plant sesquiterpene synthases

Sesquiterpene synthases (STSs) catalyze the formation of a large class of plant volatiles called sesquiterpenes. While thousands of putative STS sequences from diverse plant species are available, only a small number of them have been functionally characterized. Sequence identity-based screening for desired enzymes, often used in biotechnological applications, is difficult to apply here as STS sequence similarity is strongly affected by species. This calls for more sophisticated computational methods for functionality prediction. We investigate the specificity of precursor cation formation in these elusive enzymes. By inspecting multi-product STSs, we demonstrate that STSs have a strong selectivity towards one precursor cation. We use a machine learning approach combining sequence and structure information to accurately predict precursor cation specificity for STSs across all plant species. We combine this with a co-evolutionary analysis on the wealth of uncharacterized putative STS sequences, to pinpoint residues and distant functional contacts influencing cation formation and reaction pathway selection. These structural factors can be used to predict and engineer enzymes with specific functions, as we demonstrate by predicting and characterizing two novel STSs from Citrus bergamia. Author summaryPredicting enzyme function is a popular problem in the bioinformatics field that grows more pressing with the increase in protein sequences, and more attainable with the increase in experimentally characterized enzymes. Terpenes and terpenoids form the largest classes of natural products and find use in many drugs, flavouring agents, and perfumes. Terpene synthases catalyze the biosynthesis of terpenes via multiple cyclizations and carbocation rearrangements, generating a vast array of product skeletons. In this work, we present a three-pronged computational approach to predict carbocation specificity in sesquiterpene synthases, a subset of terpene synthases with one of the highest diversities of products. Using homology modelling, machine learning and co-evolutionary analysis, our approach combines sparse structural data, large amounts of uncharacterized sequence data, and the current set of experimentally characterized enzymes to provide insight into residues and structural regions that likely play a role in determining product specifcity. Similar techniques can be repurposed for function prediction and enzyme engineering in many other classes of enzymes.

bioinformatics↗

coronaSPAdes: from biosynthetic gene clusters to coronaviral assemblies

MotivationThe COVID-19 pandemic has ignited a broad scientific interest in viral research in general and coronavirus research in particular. The identification and characterization of viral species in natural reservoirs typically involves de novo assembly. However, existing genome, metagenome and transcriptome assemblers often are not able to assemble many viruses (including coronaviruses) into a single contig. Coverage variation between datasets and within dataset, presence of close strains, splice variants and contamination set a high bar for assemblers to process viral datasets with diverse properties. ResultsWe developed coronaSPAdes, a novel assembler for RNA viral species recovery in general and coronaviruses in particular. coronaSPAdes leverages the knowledge about viral genome structures to improve assembly extending ideas initially implemented in biosyntheticSPAdes. We have shown that coronaSPAdes outperforms existing SPAdes modes and other popular short-read metagenome and viral assemblers in the recovery of full-length RNA viral genomes. AvailabilitycoronaSPAdes version used in this article is a part of SPAdes 3.15 release and is freely available at http://cab.spbu.ru/software/spades. Contacta.korobeynikov@spbu.ru Supplementary informationSupplementary data are available at Bioinformatics

bioinformatics↗

RNAlign2D, a novel RNA structural alignment tool based on pseudo-amino acid substitution matrix

BackgroundThe functions of RNA molecules are mainly determined by their secondary structures. These functions can also be predicted using bioinformatic tools that enable the alignment of multiple RNAs to determine functional domains and/or classify RNA molecules into RNA families. However, the existing multiple RNA alignment tools, which use structural information, are slow in aligning long molecules and/or a large number of molecules. Therefore, a more rapid tool for multiple RNA alignment may improve the classification of known RNAs and help to reveal the functions of newly discovered RNAs. ResultsHere, we introduce an extremely fast Python-based tool called RNAlign2D. It converts RNA sequences to pseudo-amino acid sequences, which incorporate structural information, and uses a customizable scoring matrix to align these RNA molecules via the multiple protein sequence alignment tool MUSCLE. ConclusionsRNAlign2D produces accurate RNA alignments in a very short time. The pseudo-amino acid substitution matrix approach utilized in RNAlign2D is applicable for virtually all protein aligners.

bioinformatics↗

scREAD: A single-cell RNA-Seq database for Alzheimer's Disease

SummaryAlzheimers disease (AD) is a progressive neurodegenerative disorder of the brain and the most common form of dementia among the elderly. The single-cell RNA-sequencing (scRNA-Seq) and single-nucleus RNA-sequencing (snRNA-Seq) techniques are extremely useful for dissecting the function/dysfunction of highly heterogeneous cells in the brain at the single-cell level, and the corresponding data analyses can significantly improve our understanding of why particular cells are vulnerable in AD. We developed an integrated database named scREAD (single-cell RNA-Seq database for Alzheimers Disease), which is the first database dedicated to the management of all the existing scRNA-Seq and snRNA-Seq datasets from human postmortem brain tissue with AD and mouse models with AD pathology. scREAD provides comprehensive analysis results for 55 datasets from eight brain regions, including control atlas construction, cell type prediction, identification of differentially expressed genes, and identification of cell-type-specific regulons. Availability and ImplementationscREAD is a one-stop and user-friendly interface and freely available at https://bmbls.bmi.osumc.edu/scread/. The backend workflow can be downloaded from https://github.com/OSU-BMBL/scread/tree/master/script, to enable more discovery-driven analyses. Contactqin.ma@osumc.edu or hongjun.fu@osumc.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database

Full automation of gene prediction has become an important bioinformatics task since the advent of next generation sequencing. The eukaryotic genome annotation pipeline BRAKER1 had combined self-training GeneMark-ET with AUGUSTUS to generate genes coordinates with support of transcriptomic data. Here, we introduce BRAKER2, a pipeline with GeneMark-EP+ and AUGUSTUS externally supported by cross-species protein sequences aligned to the genome. Among the challenges addressed in the development of the new pipeline was generation of reliable hints to the locations of protein-coding exon boundaries from likely homologous but evolutionarily distant proteins. Under equal conditions, the gene prediction accuracy of BRAKER2 was shown to be higher than the one of MAKER2, yet another genome annotation pipeline. Also, in comparison with BRAKER1 supported by a large volume of transcript data, BRAKER2 could produce a better gene prediction accuracy if the evolutionary distances to the reference species in the protein database were rather small. All over, our tests demonstrated that fully automatic BRAKER2 is a fast and accurate method for structural annotation of novel eukaryotic genomes.

bioinformatics↗

Bio-informatic Analysis of Missense Single Nucleotide Polymorphisms (SNPs) in Human CD38 Gene Associated with B-Chronic Lymphocytic Leukemia

BackgroundCLL: Chronic lymphocytic leukemia is a chronic type of haematological malignancies that evoked from lymph proliferative origin of bone marrow and secondary lymphoid tissue, resultant in proliferation and progressive accumulation of distinct monoclonal CD5 /CD19 /CD23 B lymphocytes in the bone marrow, peripheral blood, and lymphatic organs. CD38 is a multifunctional ecto-enzyme, known to be a direct contributor in pathogenesis of CLL by poorly understood mechanism. Even though, it highly expressed in CLL. At specific position of CD38 gene sequence, substitution of single nucleotide may result in change in amino acid that ends by consequent alteration of protein structure. AimTo study CD38 polymorphism and to predict its effect on structure and subsequently function of CD38 molecule. Methodology and ResultThe bioinformatic analysis of CD38 gene had been carried out by using several soft wares. Functional analysis by SIFT, Polyphen2, and PROVEAN reveled 12 deleterious SNPs. These SNPs were further analyzed by SNAP2, SNP@GO. PMut, STRING and other soft wars. Furthermore, Stability analysis was done using I-Mutant and MUpro software where seven SNPs were found to decrease the stability of the protein by I-Mutant, while two SNPs increase it. At the same time, eight SNPs were found to decrease the stability by Mupro software while only one SNP is predicted to increase it. Finally, Physiochemical analysis was done using Project Hope. ConclusionIn summary, CD38 genotype seems to have twelve SNP that possibly will result in deleterious effect on Protein Structure. This genetic variation eventually will lead to alteration in potential molecule functions. Which effect the progression of CLL By the end.

bioinformatics↗

MACMIC Reveals Dual Role of CTCF in Epigenetic Regulation of Cell Identity Genes

Numerous studies of relationship between epigenomic features have focused on their strong correlation across the genome, likely because such relationship can be easily identified by many established methods for correlation analysis. However, two features with little correlation may still colocalize at many genomic sites to implement important functions. There is no bioinformatic tool for researchers to specifically identify such feature pair. Here, we develop a method to identify feature pair in which two features have maximal colocalization but minimal correlation (MACMIC) across the genome. By MACMIC analysis of 3,385 feature pairs in 15 cell types, we reveal a dual role of CTCF in epigenetic regulation of cell identity genes. Although super-enhancers are associated with activation of target genes, only a subset of super-enhancers colocalized with CTCF regulate cell identity genes. At super-enhancers colocalized with CTCF, the CTCF is required for the active marker H3K27ac in cell type requiring the activation, and also required for the repressive marker H3K27me3 in other cell types requiring the repression. Our work demonstrates the biological utility of the MACMIC analysis and reveals a key role for CTCF in epigenetic regulation of cell identity.

bioinformatics↗

SEMplMe: A tool for integrating DNA methylation effects in transcription factor binding affinity predictions

MotivationAberrant DNA methylation in transcription factor binding sites has been shown to lead to anomalous gene regulation that is strongly associated with human disease. However, the majority of methylation-sensitive positions within transcription factor binding sites remain unknown. Here we introduce SEMplMe, a computational tool to generate predictions of the effect of methylation on transcription factor binding strength in every position within a transcription factors motif. ResultsSEMplMe uses ChIP-seq and whole genome bisulfite sequencing to predict effects of methylation within binding sites. SEMplMe validates known methylation sensitive and insensitive positions within a binding motif, identifies cell type specific transcription factor binding driven by methylation, and outperforms SELEX-based predictions for CTCF. These predictions can be used to identify aberrant sites of DNA methylation contributing to human disease. Availability and ImplementationSEMplMe is available from https://github.com/Boyle-Lab/SEMplMe. Contactapboyle@umich.edu Supplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics↗

SARS-CoV-2 ORF9c Is a Membrane-Associated Protein that Suppresses Antiviral Responses in Cells

Disrupted antiviral immune responses are associated with severe COVID-19, the disease caused by SAR-CoV-2. Here, we show that the 73-amino-acid protein encoded by ORF9c of the viral genome contains a putative transmembrane domain, interacts with membrane proteins in multiple cellular compartments, and impairs antiviral processes in a lung epithelial cell line. Proteomic, interactome, and transcriptomic analyses, combined with bioinformatic analysis, revealed that expression of only this highly unstable small viral protein impaired interferon signaling, antigen presentation, and complement signaling, while inducing IL-6 signaling. Furthermore, we showed that interfering with ORF9c degradation by either proteasome inhibition or inhibition of the ATPase VCP blunted the effects of ORF9c. Our study indicated that ORF9c enables immune evasion and coordinates cellular changes essential for the SARS-CoV-2 life cycle. One-sentence summarySARS-CoV-2 ORF9c is the first human coronavirus protein localized to membrane, suppressing antiviral response, resembling full viral infection.

bioinformatics↗

MotiMul: A significant discriminative sequence motif discovery algorithm with multiple testing correction

Sequence motifs play essential roles in intermolecular interactions such as DNA-protein interactions. The discovery of novel sequence motifs is therefore crucial for revealing gene functions. Various bioinformatics tools have been developed for finding sequence motifs, but until now there has been no software based on statistical hypothesis testing with statistically sound multiple testing correction. Existing software therefore could not control for the type-1 error rates. This is because, in the sequence motif discovery problem, conventional multiple testing correction methods produce very low statistical power due to overly-strict correction. We developed MotiMul, which comprehensively finds significant sequence motifs using statistically sound multiple testing correction. Our key idea is the application of Tarones correction, which improves the statistical power of the hypothesis test by ignoring hypotheses that never become statistically significant. For the efficient enumeration of the significant sequence motifs, we integrated a variant of the PrefixSpan algorithm with Tarones correction. Simulation and empirical dataset analysis showed that MotiMul is a powerful method for finding biologically meaningful sequence motifs. The source code of MotiMul is freely available at https://github.com/ko-ichimo-ri/MotiMul.

bioinformatics↗

Training Infrastructure as a Service

BackgroundHands-on training, whether it is in Bioinformatics or other scientific domains, requires significant resources and knowledge to setup and run. Trainers must have access to infrastructure that can support the sudden spike in usage, with classes of 30 or more trainees simultaneously running resource intensive tools. For efficient classes, the jobs must run quickly, without queuing delays, lest they disrupt the timetable set out for the class. Often times this is achieved via running on a private server where there is no contention for the queue, and therefore no or minimal waiting time. However, this requires the teacher or trainer to have the technical knowledge to manage compute infrastructure, in addition to their didactic responsibilities. This presents significant burdens to potential training events, in terms of infrastructure cost, person-hours of preparation, technical knowledge, and available staff to manage such events. FindingsGalaxy Europe has developed Training Infrastructure as a Service (TIaaS) which we provide to the scientific commnuity as a service built on top of the Galaxy Platform. Training event organisers request a training and Galaxy administrators can allocate private queues specifically for the training. Trainees are transparently placed in a private queue where their jobs run without delay. Trainers access the dashboard of the TIaaS Service and can remotely follow the progress of their trainees without in-person interactions. ConclusionsTIaaS on Galaxy Europe provides reusable and fast infrastructure for Galaxy training. The instructor dashboard provides visibility into class progress, making in-person trainings more efficient and remote training possible. In the past 24 months, > 110 trainings with over 3000 trainees have used this infrastructure for training, across scientific domains, all enjoying the accessibility and reproducibility of Galaxy for training the next generation of bioinformaticians. TIaaS itself is an extension to Galaxy which can be deployed by any Galaxy administrator to provide similar benefits for their users. https://galaxyproject.eu/tiaas

bioinformatics↗

TaxonTableTools - A comprehensive, platform-independent graphical user interface software to explore and visualise DNA metabarcoding data

DNA metabarcoding is increasingly used in research and application to assess biodiversity. Powerful analysis software exists to process raw data. However, when it comes to the translation of sequence read data into biological information many end users with limited bioinformatic expertise struggle with the downstream analysis and explore data only to a minor extent. Thus, there is a growing need for easy-to-use, graphical user interface (GUI) analysis software to analyse and visualise DNA metabarcoding data. We here present TaxonTableTools (TTT), a new platform independent GUI software that aims to fill this gap by providing simple and reproducible analysis and visualisation workflows. TTT uses a so-called "TaXon table" as input. This format can easily be generated within TTT from two input files: a read table and a taxonomy table that can be obtained by various published metabarcoding pipelines. TTT analysis and visualisation modules include e.g. Venn diagrams to compare taxon overlap among replicates, samples or among different analysis methods. It analyses and visualises basic statistics such as read proportion per taxon as well as more sophisticated visualisation such as interactive Krona charts for taxonomic data exploration. Various ecological analyses such as alpha or beta diversity estimates, and rarefaction analysis ordination plots can be produced directly. Data can be explored also in formats required by traditional taxonomy-based analyses of regulatory bioassessment programs. TTT comes with a manual and tutorial, is free and publicly available through GitHub (https://github.com/TillMacher/TaxonTableTools) and the Python package index (https://pypi.org/project/taxontabletools/).

bioinformatics↗

AI aided design of epitope-based vaccine for the induction of cellular immune responses against SARS-CoV-2

The heavy burden imposed by the COVID-19 pandemic on our society triggered the race towards the development of therapies or preventive strategies. Among these, antibodies and vaccines are particularly attractive because of their high specificity, low probability of drug-drug interaction, and potentially long-standing protective effects. While the threat at hand justifies the pace of research, the implementation of therapeutic strategies cannot be exempted from safety considerations. There are several potential adverse events reported after the vaccination or antibody therapy, but two are of utmost importance: antibody-dependent enhancement (ADE) and cytokine storm syndrome (CSS). On the other hand, the depletion or exhaustion of T-cells has been reported to be associated with worse prognosis in COVID-19 patients. This observation suggests a potential role of vaccines eliciting cellular immunity, which might simultaneously limit the risk of ADE and CSS. Such risk was proposed to be associated with FcR-induced activation of proinflammatory macrophages (M1) by Fu et al. 2020 and Iwasaki et al. 2020. All aspects of the newly developed vaccine (including the route of administration, delivery system, and adjuvant selection) may affect its effectiveness and safety. In this work we use a novel in silico approach (based on AI and bioinformatics methods) developed to support the design of epitope-based vaccines. We evaluated the capabilities of our method for predicting the immunogenicity of epitopes. Next, the results of our approach were compared with other vaccine-design strategies reported in the literature. The risk of immuno-toxicity was also assessed. The analysis of epitope conservation among other Coronaviridae was carried out in order to facilitate the selection of peptides shared across different SARS-CoV-2 strains and which might be conserved in emerging zootic coronavirus strains. Finally, the potential applicability of the selected epitopes for the development of a vaccine eliciting cellular immunity for COVID-19 was discussed, highlighting the benefits and challenges of such an approach.

bioinformatics↗