bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Deduplication Improves Cost-Efficiency and Yields of De novo Assembly and Binning of Shot-Gun Metagenomes in Microbiome Research

Metagenomics has in the last decade greatly revolutionized the study of microbial communities. However, the presence of artificial duplicate reads mainly raised from the preparation of metagenomic DNA sequencing library and their impacts on metagenomic assembly and binning have never brought to the attention. Here, we explicitly investigated the effects of duplicate reads on metagenomic assembly and binning, based on analyses of four groups of representative metagenomes with distinct microbiome complexity. Our results showed that deduplication considerably increased the binning yields (by 3.5% to 80%) for most of the metagenomic datasets examined thanks to improved contig length and coverage profiling of metagenome-assembled contigs. Specifically, 411 versus 397, 331 versus 317, 104 versus 88 and 9 versus 5 metagenome-assembled genomes (MAGs) were recovered from MEGAHIT assemblies of bioreactor sludge, surface water, lake sediment, and forest soil metagenomes, respectively. Noticeably, deduplication reduced the computational costs of metagenomic assembly including elapsed time (by 9.0% to 29.9%) and maximum memory requirement (by 4.3% to 37.1%). Collectively, it is recommended to remove duplicate reads in metagenomic data before assembly and binning analyses, particularly for complex environmental samples, such as forest soils examined in this study. ImportanceDuplicated reads are usually considered as technical artefacts. Their presence in metagenomes would theoretically not only introduce bias in the quantitative analysis, but also result in mistakes in coverage profile, leading to negative effects or even failures on metagenomic assembly and binning, as the widely used metagenome assemblers and binners all need coverage information for graph partitioning and assembly binning, respectively. However, this issue was seldomly noticed and its impacts on the downstream key bioinformatic procedures (e.g., assembly and binning) still remained unclear. In this study, we comprehensively evaluated for the first time the impacts of duplicate reads on de novo assembly and binning of real metagenomic datasets by comparing assembly quality, binning yields and the requirements of computational resources with and without the removal of duplicate reads. It was revealed that deduplication considerably increased the binning yields and significantly reduced the computational costs including elapsed time and maximum memory requirement. The results provide empirical reference for more cost-efficient metagenomic analyses in microbiome research.

bioinformatics↗

In silico comparative RNA-seq analysis reveals varietal-specific intergenic small open reading frames in Cucumis sativus L.

Small open reading frames (sORFs) have been reported to play important roles in growth, regulation of morphogenesis, and abiotic stress responses in various plant species. However, their sequences and functions remain poorly understood in many plant species including Cucumis sativus. Cucumis sativus (commonly known as cucumber) is Asias fourth most important vegetable and the second most important crop in Western Europe. The breeding of climate-resilient cucumbers is of great importance to ensure their sustainability under extreme climate conditions. In this study, we aim to isolate the intergenic sORFs from C. sativus var. hardwickii genome and determine their sequence diversity and expression profiles in C. sativus var. hardwickii and different cultivars of C. sativus var. sativus using bioinformatics tools. We identified a total of 50,191 coding sORFs with coding potential (coding sORFs) from C. sativus var. hardwickii genome. In addition, 1,311 transcribed sORFs were detected in RNA-seq datasets of C. sativus var. hardwickii and shared homology to sequences deposited in the cucumber EST database, and among these, 91 transcribed sORFs with translation potential were detected. A total of 629 high-confident C. sativus-specific sORFs were identified in both varieties. Varietal-specific transcribed sORFs were also identified in C. sativus var. hardwickii (87) and C. sativus var. sativus (2,906). Furthermore, cultivar- and tissue-specific transcribed sORFs were identified in different cultivars and tissue samples. The findings of this study provide insight into sequence diversity and expression patterns of sORFs in C. sativus, which could help in developing climate-resilient cucumbers.

bioinformatics↗

Entropy predicts fuzzy-seed sensitivity

In sequence similarity search applications such as read mapping, it is desired that seeds match between a read and reference in regions with mutations or read errors (seed sensitivity). K-mers are likely the most well-known and used seed construct in bioinformatics, and many studies on, e.g., spaced k-mers aim to improve sensitivity over k-mers. Spaced k-mers are highly sensitive when substitutions largely dominate the mutation rate but quickly deteriorate when indels are present. Recently, we developed a pseudo-random seeding construct, strobemers, which were empirically demonstrated to have high sensitivity also at high indel rates. However, the study lacked a deeper understanding of why. In this study, we demonstrate that a seeds entropy (randomness) is a good predictor for seed sensitivity. We propose a model to estimate the entropy of a seed and find that seeds with high entropy, according to our model, in most cases have high match sensitivity. We also present three new strobemer seed constructs, mixedstrobes, altstrobes, and multistrobes. We use both simulated and biological data to demonstrate that our new seed constructs improve sequence-matching sensitivity to other strobemers. We implement strobemers into minimap2 and observe slightly faster alignment time and higher accuracy than using k-mers at various error rates. Our discovered seed randomness-sensitivity relationship explains why some seeds perform better than others, and the relationship provides a framework for designing even more sensitive seeds. In addition, we show that the three new seed constructs are practically useful. Finally, in cases where our entropy model does not predict the observed sensitivity well, we explain why and how to improve the model in future work.

bioinformatics↗

Fast and Accurate Prediction of Intrinsically Disordered Protein by Protein Language Model

MotivationIntrinsically disordered proteins (IDPs) play a vital role in various biological processes and have attracted increasing attention in the last decades. Predicting IDPs from primary structures of proteins provides a very useful tool for protein analysis. However, most of the existing prediction methods heavily rely on multiple sequence alignments (MSAs) of homologous sequences which are formed by evolution over billions of years. Obtaining such information requires searching against the whole protein databases to find similar sequences and since this process becomes increasingly time-consuming, especially in large-scale practical applications, the alternative method is needed. ResultsIn this paper, we proposed a novel IDP prediction method named IDP-PLM, based on the protein language model (PLM). The method does not rely on MSAs or MSA-based profiles but leverages only the protein sequences, thereby achieving state-of-the-art performance even compared with predictors using protein profiles. The proposed IDP-PLM is composed of stacked predictors designed for several different protein-related tasks: secondary structure prediction, linker prediction, and binding predictions. In addition, predictors for the single task also achieved the highest accuracy. All these are based on PLMs thus making IDP-PLM not rely on MSA-based profiles. The ablation study reveals that all these stacked predictors contribute positively to the IDP prediction performance of IDP-PLM. AvailabilityThe method is available at http://github.com/xu-shi-jie. Contactakira.onoda@ees.hokudai.ac.jp Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

aRgus: multilevel visualization of non-synonymous single nucleotide variants & advanced pathogenicity score modeling for genetic vulnerability assessment

The widespread use of high-throughput sequencing techniques is leading to a rapidly increasing number of disease-associated variants of unknown significance and candidate genes. Integration of knowledge concerning their genetic, protein as well as functional and conservational aspects is necessary for an exhaustive assessment of their relevance and for prioritization of further clinical and functional studies investigating their role in human disease. In order to collect the necessary information, a multitude of different databases has to be accessed and data extraction from the original sources commonly is not user-friendly and requires advanced bioinformatics skills. This leads to a decreased data accessibility for a relevant number of potential users such as clinicians, geneticist, and clinical researchers. Here, we present aRgus (https://argus.urz.uni-heidelberg.de/), a standalone webtool for simple extraction and intuitive visualization of multi-layered gene, protein, variant, and variant effect prediction data. aRgus provides interactive exploitation of these data within seconds for any known gene of the human genome. In contrast to existing online platforms for compilation of variant data, aRgus complements visualization of chromosomal exon-intron structure and protein domain annotation with ClinVar and gnomAD variant distributions as well as position-specific variant effect prediction score modeling. aRgus thereby enables timely assessment of protein regions vulnerable to variation with single amino acid resolution and provides numerous applications in variant and protein domain interpretation as well as in the design of in vitro experiments.

bioinformatics↗

Pathway-informed deep learning model for survivalanalysis and pathological classification of gliomas

Online assessment of tumor characteristics during surgery is important and has the potential to establish an intraoperative surgeon feedback mechanism. With the availability of such feedback, surgeons could decide to be more liberal or conservative regarding the resection of the tumor. While there are methods to perform metabolomics-based online tumor pathology prediction, their model complexity and, in turn, the predictive performance is limited by the small dataset sizes. Furthermore, the information conveyed by the feedback provided on the tumor tissue could be improved both in terms of content and accuracy. In this study, we propose a metabolic pathway-informed deep learning model, PiDeeL, to perform survival analysis and pathology assessment based on metabolite concentrations. We show that incorporating pathway information into the model architecture substantially reduces parameter complexity and achieves better survival analysis and pathological classification performance. With these design decisions, we show that PiDeeL improves tumor pathology prediction performance of the state-of-the-art in terms of the Area Under the ROC Curve (AUC-ROC) by 3.38% and the Area Under the Precision-Recall Curve (AUC-PR) by 4.06%. Similarly, with respect to the time-dependent concordance index (c-index), we observe that PiDeeL achieves better survival analysis performance (improvement up to 4.3%) when compared to the state-of-the-art. Moreover, we show that importance analyses performed on input metabolite features as well as pathway-specific hidden-layer neurons of PiDeeL provide insights into tumor metabolism. We foresee that the use of this model in the surgery room will help surgeons adjust the surgery plan on the fly and will result in better prognosis estimates tailored to surgical procedures. AvailabilityThe code is released at https://github.com/ciceklab/PiDeeL. The data used in this study is released at https://zenodo.org/record/7228791. Contactcicek@cs.bilkent.edu.tr Supplementary informationSupplementary data are available at Briefings in Bioinformatics online.

bioinformatics↗

ViReaDB: A user-friendly database for compactly storing viral sequence data and rapidly computing consensus genome sequences

MotivationIn viral molecular epidemiology, reconstruction of consensus genomes from sequence data is critical for tracking mutations and variants of concern. However, storage of the raw sequence data can become prohibitively large, and computing consensus genome from sequence data can be slow and requires bioinformatics expertise. ResultsViReaDB is a user-friendly database system for compactly storing viral sequence data and rapidly computing consensus genome sequences. From a dataset of 1 million trimmed mapped SARS-CoV-2 reads, it is able to compute the base counts and the consensus genome in 16 minutes, store the reads alongside the base counts and consensus in 50 MB, and optionally store just the base counts and consensus (without the reads) in 300 KB. AvailabilityViReaDB is freely available on PyPI (https://pypi.org/project/vireadb) and on GitHub (https://github.com/niemasd/ViReaDB) as an open-source Python software project. Contactniema@ucsd.edu

bioinformatics↗

Technical report on best practices for hybrid and long read de novo assembly of bacterial genomes utilizing Illumina and Oxford Nanopore Technologies reads

The emergence of commercial long read sequencing technologies in the 2010s and the concomitant development of new bioinformatics tools bears the potential of de novo genome assemblies of unprecedented contiguity and quality. However, until today these novel technologies suffer from high rates of sequencing errors. These may be overcome by using long and short reads in combination, in so called hybrid approaches, or by increasing the through-put and thereby the coverage of sequencing runs. In particular the latter will thereby increase the cost of the assembly inevitably. Herein, to-date long read and hybrid assemblers were tested on real whole genome sequencing Illumina and Oxford Nanopore Technologies read data sets and sub samples of these in order to elaborate a best practice for de novo assembly. The findings suggest that although long reads alone can be used to reconstruct complete and contiguous genomes, in particular the single-nucleotide and indel error rate remains high compared to hybrid approaches and that this can impact downstream applications such as variation discovery and gene prediction negatively.

bioinformatics↗

Aligning Distant Sequences to Graphs using Long Seed Sketches

Sequence-to-graph alignment is an important step in applications such as variant genotyping, read error correction and genome assembly. When a query sequence requires a substantial number of edits to align, approximate alignment tools that follow the seed-and-extend approach require shorter seeds to get any matches. However, in large graphs with high variation, relying on a shorter seed length leads to an exponential increase in spurious matches. We propose a novel seeding approach relying on long inexact matches instead of short exact matches. We demonstrate experimentally that our approach achieves a better time-accuracy trade-off in settings with up to a 25% mutation rate. We achieve this by sketching a subset of graph nodes and storing them in a K-nearest neighbor index. While sketches are more robust to indels, finding the nearest neighbor of a sketch in a high-dimensional space is more computationally challenging than finding exact seeds. We demonstrate that if we store sketch vectors in a K-nearest neighbor index, we can circumvent the curse of dimensionality. Our long sketch-based seed scheme contrasts existing approaches and highlights the important role that tensor sketching can play in bioinformatics applications. Our proposed seeding method and implementation have several advantages: i) We empirically show that our method is efficient and scales to graphs with 1 billion nodes, with time and memory requirements for preprocessing growing linearly with graph size and query time growing quasi-logarithmically with query length. ii) For queries with an edit distance of 25% relative to their length, on the 1 billion node graph, longer sketch-based seeds yield a 4x increase in recall compared to exact seeds. iii) Conceptually, our seeder can be incorporated into other aligners, proposing a novel direction for sequence-to-graph alignment. The implementation is available at: https://github.com/ratschlab/tensor-sketch-alignment.

bioinformatics↗

Statistically Consistent Rooting of Species Trees under the Multi-Species Coalescent Model

Rooted species trees are used in several downstream applications of phylogenetics. Most species tree estimation methods produce unrooted trees and additional methods are then used to root these unrooted trees. Recently, Quintet Rooting (QR) (Tabatabaee et al., ISMB and Bioinformatics 2022), a polynomial-time method for rooting an unrooted species tree given unrooted gene trees under the multispecies coalescent, was introduced. QR, which is based on a proof of identifiability of rooted 5-taxon trees in the presence of incomplete lineage sorting, was shown to have good accuracy, improving over other methods for rooting species trees when incomplete lineage sorting was the only cause of gene tree discordance, except when gene tree estimation error was very high. However, the statistical consistency of QR was left as an open question. Here, we present QR-STAR, a polynomial-time variant of QR that has an additional step for determining the rooted shape of each quintet tree. We prove that QR-STAR is statistically consistent under the multispecies coalescent model. Our simulation study under a variety of model conditions shows that QR-STAR matches or improves on the accuracy of QR. QR-STAR is available in open source form at https://github.com/ytabatabaee/Quintet-Rooting.

bioinformatics↗

Spectrum preserving tilings enable sparse and modular reference indexing

The reference indexing problem for k-mers is to pre-process a collection of reference genomic sequences[R] so that the position of all occurrences of any queried k-mer can be rapidly identified. An efficient and scalable solution to this problem is fundamental for many tasks in bioinformatics. In this work, we introduce the spectrum preserving tiling (SPT), a general representation of[R] that specifies how a set of tiles repeatedly occur to spell out the constituent reference sequences in[R] . By encoding the order and positions where tiles occur, SPTs enable the implementation and analysis of a general class of modular indexes. An index over an SPT decomposes the reference indexing problem for k-mers into: (1) a k-mer-to-tile mapping; and (2) a tile-to-occurrence mapping. Recently introduced work to construct and compactly index k-mer sets can be used to efficiently implement the k-mer-to-tile mapping. However, implementing the tile-to-occurrence mapping remains prohibitively costly in terms of space. As reference collections become large, the space requirements of the tile-to-occurrence mapping dominates that of the k-mer-to-tile mapping since the former depends on the amount of total sequence while the latter depends on the number of unique k-mers in[R] . To address this, we introduce a class of sampling schemes for SPTs that trade off speed to reduce the size of the tile-to-reference mapping. We implement a practical index with these sampling schemes in the tool pufferfish2. When indexing over 30,000 bacterial genomes, pufferfish2 reduces the size of the tile-to-occurrence mapping from 86.3GB to 34.6GB while incurring only a 3.6x slowdown when querying k-mers from a sequenced readset. Supplementary materialsSections S.1 to S.8 available online at https://doi.org/10.5281/zenodo.7504717 Availabilitypufferfish2 is implemented in Rust and available at https://github.com/COMBINE-lab/pufferfish2.

bioinformatics↗

Toblerone: detecting exon deletion events in cancer using RNA-seq

Cancer is driven by mutations of the genome that can result in the activation of oncogenes or repression of tumour suppressor genes. In acute lymphoblastic leukemia (ALL) focal deletions in IKAROS family zinc finger 1 (IKZF1) result in the loss of zinc-finger DNA-binding domains and a dominant negative isoform that is associated with higher rates of relapse and poorer patient outcomes. Clinically, the presence of IKZF1 deletions informs prognosis and treatment options. In this work we developed a method for detecting exon deletions in genes using RNA-seq with application to IKZF1. We developed a pipeline that first uses a custom transcriptome reference consisting of transcripts with exon deletions. Next, RNA-seq reads are mapped using a pseudoalignment algorithm to identify reads that uniquely support deletions. These are then evaluated for evidence of the deletion with respect to gene expression and other samples. We applied the algorithm, named Toblerone, to a cohort of 99 B-ALL paediatric samples including validated IKZF1 deletions. Furthermore, we developed a graphical desktop app for non-bioinformatics users that can quickly and easily identify and report deletions in IKZF1 from RNA-seq data with informative graphical outputs.

bioinformatics↗

Sashimi.py: a flexible toolkit for combinatorial analysis of genomic data

Simultaneously visualizing how isoform expression, protein-DNA/RNA interactions, accessibility, and architecture of chromatin differs across condition and cell types could inform our understanding on regulatory mechanisms and functional consequences of alternative splicing. However, the existing versions of tools generating sashimi plots remain inflexible, complicated, and user-unfriendly for integrating data sources from multiple bioinformatic formats or various genomics assays. Thus, a more scalable visualization tool is necessary to broaden the scope of sashimi plots. Here, we introduce sashimi.py, a Python package for generating publication-quality visualization by a programmable and interactive web-based approach. Sashimi.py is a platform to visually interpret genomic data from a large variety of data sources including single-cell RNA-seq, protein-DNA/RNA interactions, long-reads sequencing data, and Hi-C data without any preprocessing, and also offers a broad degree of flexibility for formats of output files that satisfy the requirements of major journals. The Sashimi.py package is an open-source software which is freely available on Bioconda (https://anaconda.org/bioconda/sashimi-py), Docker, PyPI (https://pypi.org/project/sashimi.py/) and GitHub (https://github.com/ygidtu/sashimi.py), and a built-in web server for local deployment is also provided.

bioinformatics↗

Machine-learning based detection of adventitious microbes in T-cell therapy cultures using long read sequencing

Assuring that cell therapy products are safe before releasing them for use in patients is critical. Currently, compendial sterility testing for bacteria and fungi can take 7-14 days. The goal of this work was to develop a rapid untargeted approach for the sensitive detection of microbial contaminants at low abundance from low volume samples during the manufacturing process of cell therapies. We developed a long-read sequencing methodology using Oxford Nanopore Technologies MinION platform with 16S and 18S amplicon sequencing to detect USP<71> organisms and other microbial species. Reads are classified metagenomically to predict the microbial species. We used an extreme gradient boosting machine learning algorithm (XGBoost) to first assess if a sample is contaminated and second, determine whether the predicted contaminant is correctly classified or misclassified. The model was used to make a final decision on the sterility status of the input sample. An optimised experimental and bioinformatics pipeline starting from spiked species through to sequenced reads allowed for the detection of microbial samples at 10 CFU / mL using metagenomic classification. Machine learning can be coupled with long read sequencing to detect and identify sample sterility status and microbial species present in T-cell cultures, including the USP<71> organisms to 10 CFU / mL. ImportanceThis research presents a novel method for rapidly and accurately detecting microbial contaminants in cell therapy products, which is essential for ensuring patient safety. Traditional testing methods are time-consuming, taking 7-14 days, while our approach can significantly reduce this time. By combining advanced long read Nanopore sequencing techniques and machine learning, we can effectively identify the presence and types of microbial contaminants at low abundance levels. This breakthrough has the potential to improve the safety and efficiency of cell therapy manufacturing, leading to better patient outcomes and a more streamlined production process.

bioinformatics↗

Improving the Quality of Co-evolution Intermolecular Contact Prediction with DisVis

The steep rise in available protein sequences and structures has paved the way for bioinformatics approaches to predict residue-residue interactions in protein complexes. Multiple sequence alignments are commonly used in intermolecular contact predictions to identify co-evolving residues. These contacts, however, often include false positives (FPs), which may impair their use to predict three dimensional structures of biomolecular complexes and affect the accuracy of the generated models. Previously, we have developed DisVis to identify false positive data in mass spectrometry cross-linking data. DisVis allows to assess the accessible interaction space between two proteins consistent with a set of distance restraints. Here, we investigate if a similar approach could be applied to co-evolution predicted contacts in order to improve their precision prior to using them for modelling complexes. In this work we analyze co-evolution contact predictions with DisVis in order to identify putative FPs for a set of 26 protein-protein complexes. Next, the DisVis-reranked and the original co-evolution contacts are used to model the complexes with our integrative docking software HADDOCK using different filtering scenarios. Our results show that HADDOCK is robust with respect to the precision of the predicted contacts due to the 50% random contact removal during docking and using DisVis filtering for low precision contact data. DisVis can thus have a beneficial effect on low quality data, but overall HADDOCK can accommodate FP restraints without negatively impacting the quality of the resulting models. Other more precision-sensitive docking protocols might, however, benefit from the increased precision of the predicted contacts after DisVis filtering.

bioinformatics↗

A graph clustering algorithm for detection and genotyping of structural variants from long reads

Structural variants (SV) are polymorphisms defined by their length (>50 bp). The usual types of SVs are deletions, insertions, translocations, inversions, and copy number variants. SV detection and genotyping is fundamental given the role of SVs in phenomena such as phenotypic variation and evolutionary events. Thus, methods to identify SVs using long read sequencing data have been recently developed. We present an accurate and efficient algorithm to predict SVs from long-read sequencing data. The algorithm starts collecting evidence (Signatures) of SVs from read alignments. Then, signatures are clustered based on a Euclidean graph with coordinates calculated from lengths and genomic positions. Clustering is performed by the DBSCAN algorithm, which provides the advantage of delimiting clusters with high resolution. Clusters are transformed into SVs and a Bayesian model allows to precisely genotype SVs based on their supporting evidence. This algorithm is integrated in the single sample variants detector of the Next Generation Sequencing Experience Platform (NGSEP), which facilitates the integration with other functionalities for genomics analysis. For benchmarking, our algorithm is compared against different tools using VISOR for simulation and the GIAB SV dataset for real data. For indel calls in a 20x depth Nanopore simulated dataset, the DBSCAN algorithm performed better, achieving an F-score of 98%, compared to 97.8 for Dysgu, 97.8 for SVIM, 97.7 for CuteSV, and 96.8 for Sniffles. We believe that this work makes a significant contribution to the development of bioinformatic strategies to maximize the use of long read sequencing technologies.

bioinformatics↗

Combined promoter-capture Hi-C and Hi-C analysis reveals a fine-tuned regulation of 3D chromatin architecture in colorectal cancer

Hi-C is a widely used method for profiling chromosomal interactions in the 3-dimensional context. Due to limitations on the depth of sequencing, the resolution of most Hi-C datasets is often insufficient for scoring fine-scale interactions. We therefore used promoter-capture Hi-C (PCHi-C) data for mapping these subtle interactions. From multiple colorectal cancer (CRC) studies, we combined PCHi-C with Hi-C datasets to understand the dynamics of chromosomal interactions from cis regulatory elements to topologically associated domain (TAD)-level, enabling detection of fine-scale interactions of disease-associated loci within TADs. Our integrated analyses of PCHi-C and Hi-C datasets from CRC cell lines along with histone modification landscape and transcriptome signatures highlight significant genomic structural instability and their association with tumor-suppressive transcriptional programs. Such analyses also yielded nine dysregulated genes. Transcript profiling revealed a dramatic increase in their expression in CRC cell lines as compared to NT2D1 human embryonic carcinoma cells, supporting the predictions of our bioinformatics analysis. We further report increased occupancy of activation associated histone modifications H3K27ac and H3K4me3 at the promoter regions of the targets analyzed. Our study provides deeper insights into the dynamic 3D genome organization in CRC and identification of affected genes which may serve as potential biomarkers for CRC. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=109 SRC="FIGDIR/small/515643v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@11a9fc8org.highwire.dtl.DTLVardef@f01eb4org.highwire.dtl.DTLVardef@6ff038org.highwire.dtl.DTLVardef@103fc24_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

FASSO: An AlphaFold based method to assign functional annotations by combining sequence and structure orthology

Methods to predict orthology play an important role in bioinformatics for phylogenetic analysis by identifying orthologs within or across any level of biological classification. Sequence-based reciprocal best hit approaches are commonly used in functional annotation since orthologous genes are expected to share functions. The process is limited as it relies solely on sequence data and does not consider structural information and its role in function. Previously, determining protein structure was highly time-consuming, inaccurate, and limited to the size of the protein, all of which resulted in a structural biology bottleneck. With the release of AlphaFold, there are now over 200 million predicted protein structures, including full proteomes for dozens of key organisms. The reciprocal best structural hit approach uses protein structure alignments to identify structural orthologs. We propose combining both sequence- and structure-based reciprocal best hit approaches to obtain a more accurate and complete set of orthologs across diverse species, called Functional Annotations using Sequence and Structure Orthology (FASSO). Using FASSO, we annotated orthologs between five plant species (maize, sorghum, rice, soybean, Arabidopsis) and three distance outgroups (human, budding yeast, and fission yeast). We inferred over 270,000 functional annotations across the eight proteomes including annotations for over 5,600 uncharacterized proteins. FASSO provides confidence labels on ortholog predictions and flags potential misannotations in existing proteomes. We further demonstrate the utility of the approach by exploring the annotation of the maize proteome.

bioinformatics↗