bioRxiv ScienceSearch

Biology subjects

Bernardes, J. S.

Publications and source records attributed to Bernardes, J. S..

2 recordsLinked to original sources

DAVI: a tool for clustering and visualising protein domain architectures

The characterization of protein functions is one of the main challenges in bioinformatics. Proteins are often composed of individual units termed domains, motifs that can evolve independently. The domain architecture of a given protein is the particular order and the content of its numerous domains. Some computational approaches predict the most likely domain architecture for a set of proteins. However, a few numbers of visualization tools exist, and most of them are unavailable. Here we present DAVI, an efficient and user-friendly web server for protein domain architecture clustering and visualization. DAVI accepts the output of most used domain architecture prediction tools and also produces domain architectures for a set of protein sequences. It provides a rich visualization for comparing, analyzing, and visualizing domain architectures. Availabilityhttp://genome.lcqb.upmc.fr/Domain-Architecture-Viewer

bioinformatics

Automatic generation of ground truth data for the evaluation of clonal grouping methods in B-cell populations

MotivationThe adaptive B-cell response is driven by the expansion, somatic hypermutation, and selection of B-cell clones. Their number, size and sequence diversity are essential characteristics of B-cell populations. Identifying clones in B-cell populations is central to several repertoire studies such as statistical analysis, repertoire comparisons, and clonal tracking. Several clonal grouping methods have been developed to group sequences from B-cell immune repertoires. Such methods have been principally evaluated on simulated benchmarks since experimental data containing clonally related sequences can be difficult to obtain. However, experimental data might contains multiple sources of sequence variability hampering their artificial reproduction. Therefore, the generation of high precision ground truth data that preserves real repertoire distributions is necessary to accurately evaluate clonal grouping methods. ResultsWe proposed a novel methodology to generate ground truth data sets from real repertoires. Our procedure requires V(D)J annotations to obtain the initial clones, and iteratively apply an optimisation step that moves sequences among clones to increase their cohesion and separation. We first showed that our method was able to identify clonally-related sequences in simulated repertoires with higher mutation rates, accurately. Next, we demonstrated how real benchmarks (generated by our method) constitute a challenge for clonal grouping methods, when comparing the performance of a widely used clonal grouping algorithm on several generated benchmarks. Our method can be used to generate a high number of benchmarks and contribute to construct more accurate clonal grouping tools. Availability and implementationThe source code and generated data sets are freely available at github.com/NikaAb/BCR_GTG

bioinformatics