bioRxiv Science⌕ Search

Biology subjects

Molari, M.

Publications and source records attributed to Molari, M..

4 recordsLinked to original sources

PanGraph: scalable bacterial pan-genome graph construction

The genomic diversity of microbes is commonly parameterized as single nucleotide polymorphisms relative to a reference genome of a well-characterized, but arbitrary, isolate. However, any reference genome contains only a fraction of the microbial pangenome, the total set of genes observed in a given species. Reference-based approaches are thus blind to the dynamics of the accessory genome, as well as variation within gene order and copy number. With the wide-spread usage of long-read sequencing, the number of high-quality, complete genome assemblies has increased dramatically. Traditional computational approaches towards whole-genome analysis either scale poorly with the number of genomes, or treat genomes as dissociated "bags of genes", and thus are not suited for this new era. Here, we present PanGraph, a Julia-based library and command line interface for aligning whole genomes into a graph. Each genome is represented as an undirected path along vertices, which in turn, encapsulate homologous multiple sequence alignments. The resultant data structure succinctly summarizes population-level nucleotide and structural polymorphisms and can be exported into a several common formats for either downstream analysis or immediate visualization.

bioinformatics↗

FAIR enough? A perspective on the status of nucleotide sequence data and metadata on public archives

Knowledge derived from nucleotide sequence data is increasing in importance in the life sciences, as well as decision making (mainly in biodiversity policy). Metadata standards have been established to facilitate sustainable sequence data management according to the FAIR principles (Findability, Accessibility, Interoperability, Reusability). Here, we review the status of metadata available for raw read Illumina amplicon and whole genome shotgun sequencing data derived from ecological metagenomic material that are accessible at the European Nucleotide Archive (ENA), as well as the compliance of the primary sequence data (fastq files) with data submission requirements. While overall basic metadata, such as geographic coordinates, were retrievable in 98% of the cases for this type of sequence data, interoperability was not always ensured and other (mainly conditionally) mandatory parameters were often not provided at all. Metadata standards, such as the Minimum Information about any(x) Sequence (MIxS), were only infrequently used despite a demonstrated positive impact on metadata quality. Furthermore, the sequence data itself did not meet the prescribed requirements in 31 out of 39 studies that were manually inspected. To tackle the most immediate needs to improve FAIR sequence data management, we provide a list of minimal suggestions to researchers, research institutions, funding agencies, reviewers, publishers, and databases, that we believe might have a potentially large positive impact on sequence data and metadata FAIRness, which is crucial for further research and its derived applications.

molecular biology↗

The hidden pangenome: comparative genomics reveals pervasive diversity in symbiotic and free-living sulfur-oxidizing bacteria

Sulfur-oxidizing Thioglobaceae, often referred to as SUP05 and Arctic96BD clades, are widespread and common to hydrothermal vents and oxygen minimum zones. They impact global biogeochemical cycles and exhibit a variety of host-associated and free-living lifestyles. The evolutionary driving forces that led to the versatility, adoption of multiple lifestyles and global success of this family are largely unknown. Here, we perform an in-depth comparative genomic analysis using all available and newly generated Thioglobaceae genomes. Gene content variation was common, throughout taxonomic ranks and lifestyles. We uncovered a pool of variable genes within most Thioglobaceae populations in single environmental samples and we referred to this as the hidden pangenome. The hidden pangenome is often overlooked in comparative genomic studies and our results indicate a much higher intra-specific diversity within environmental bacterial populations than previously thought. Our results show that core-community functions are different from species core genomes suggesting that core functions across populations are divided among the intra-specific members within a population. Defense mechanisms against foreign DNA and phages were enriched in symbiotic lineages, indicating an increased exchange of genetic material in symbioses. Our study suggests that genomic plasticity and frequent exchange of genetic material drives the global success of this family by increasing its evolvability in a heterogeneous environment.

microbiology↗

Quantitative modeling of the effect of antigen dosage on B-cell affinity distributions in maturating germinal centers

Affinity maturation is a complex dynamical process allowing the immune system to generate antibodies capable of recognizing antigens. We introduce a model for the evolution of the distribution of affinities across the antibody population in germinal centers. The model is amenable to detailed mathematical analysis, and gives insight on the mechanisms through which antigen availability controls the rate of maturation and the expansion of the antibody population. It is also capable, upon maximum-likelihood inference of the parameters, to reproduce accurately the distributions of affinities of IgG-secreting cells we measure in mice immunized against Tetanus Toxoid under largely varying conditions (antigen dosage, delay between injections). Both model and experiments show that the average population affinity depends non-monotonically on the antigen dosage. We show that combining quantitative modelling and statistical inference is a concrete way to investigate biological processes underlying affinity maturation (such as selection permissiveness), hardly accessible through measurements.

immunology↗