bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

Comparative genome analysis using sample-specific string detection in accurate long reads

MotivationComparative genome analysis of two or more whole-genome sequenced (WGS) samples is at the core of most applications in genomics. These include discovery of genomic differences segregating in population, case-control analysis in common diseases, and rare disorders. With the current progress of accurate long-read sequencing technologies (e.g., circular consensus sequencing from PacBio sequencers) we can dive into studying repeat regions of genome (e.g., segmental duplications) and hard-to-detect variants (e.g., complex structural variants). ResultsWe propose a novel framework for addressing the comparative genome analysis by discovery of strings that are specific to one genome ("samples-specific" strings). We have developed an accurate and efficient novel method for discovery of samples-specific strings between two groups of WGS samples. The proposed approach will give us the ability to perform comparative genome analysis without the need to map the reads and is not hindered by shortcomings of the reference genome. We show that the proposed approach is capable of accurately finding samples-specific strings representing nearly all variation (> 98%) reported across pairs or trios of WGS samples using accurate long reads (e.g., PacBio HiFi data). AvailabilityThe proposed tool is publicly available at https://github.com/Parsoa/PingPong.

bioinformatics↗

The non-genomic vitamin D pathway links β-amyloid to autophagic apoptosis in Alzheimer's disease

Vitamin D is an important hormonal molecule, which exerts genomic and non-genomic actions in maintaining brain development and adult brain health. Many epidemiological studies have associated vitamin D deficiency with Alzheimers disease (AD). Nevertheless, the underlying signaling pathway through which this occurs remains to be characterized. We were intrigued to find that although vitamin D levels are significantly low in AD patients, their hippocampal vitamin D receptor (VDR) levels are inversely increased in the cytosol of the brain cells, and colocalized with A{beta}42 plaques, gliosis and autophagosomes, suggesting that a non-genomic form of VDR is implicated in AD. Mechanistically, A{beta}42 induces the conversion of nuclear heterodimer of VDR/RXR heterodimer into a cytoplasmic VDR/p53 heterodimer. The cytosolic VDR/p53 complex mediates the A{beta}42-induced autophagic apoptosis. Reduction of p53 activity in AD mice reverses the VDR/RXR formation and rescues AD brain pathologies and cognitive impairment. In line with the impaired genomic VDR pathway, the transgenic AD mice fed a vitamin D sufficient diet exhibit lower plasma vitamin D levels since early disease phases, raising the possibility that vitamin D deficiency may actually be an early manifestation of AD. Despite the deficiency of vitamin D in AD mice, vitamin D supplementation not only has no benefit but lead to exacerbated A{beta}42 depositions and cognitive impairment. Together, these data indicate that the impaired genomic vitamin D pathway links A{beta}42 to induce autophagic apoptosis, and suggest that VDR/p53 pathway could be targeted for the treatment of AD. Significance StatementVitamin D exerts a genomic action for neuroprotection through VDR/RXR transcriptional complex. Thus, insufficient vitamin D has been linked to AD, but the signaling pathway involved remains unclear. Surprisingly, we find that the genomic action of VDR/RXR to be compromised and converted into a non-genomic VDR/p53 complex in promoting AD neurodegeneration. The cytosolic VDR/p53 complex contribute to autophagy-induced neuronal apoptosis. The VDR/RXR pathway can be a new therapeutic target for AD because targeting VDR/p53 ameliorates AD. Importantly, we provide evidence that vitamin D deficiency might be an early AD manifestation, and vitamin D supplementation exacerbates AD. This work uncovers a non-genomic VDR action in promoting AD and suggests a potential aggravating effect of vitamin D supplementation on AD.

neuroscience↗

Genome size evolution in the diverse insect order Trichoptera

BackgroundGenome size is implicated in form, function, and ecological success of a species. Two principally different mechanisms are proposed as major drivers of eukaryotic genome evolution and diversity: Polyploidy (i.e., whole genome duplication: WGD) or smaller duplication events and bursts in the activity of repetitive elements (RE). Here, we generated de novo genome assemblies of 17 caddisflies covering all major lineages of Trichoptera. Using these and previously sequenced genomes, we use caddisflies as a model for understanding genome size evolution in diverse insect lineages. ResultsWe detect a ~14-fold variation in genome size across the order Trichoptera. We find strong evidence that repetitive element (RE) expansions, particularly those of transposable elements (TEs), are important drivers of large caddisfly genome sizes. Using an innovative method to examine TEs associated with universal single copy orthologs (i.e., BUSCO genes), we find that TE expansions have a major impact on protein-coding gene regions, with TE-gene associations showing a linear relationship with increasing genome size. Intriguingly, we find that expanded genomes preferentially evolved in caddisfly clades with a higher ecological diversity (i.e., various feeding modes, diversification in variable, less stable environments). ConclusionOur findings provide a platform to test hypotheses about the potential evolutionary roles of TE activity and TE-gene associations, particularly in groups with high species, ecological, and functional diversities.

evolutionary biology↗

A widely distributed genus of soil Acidobacteria genomically enriched in biosynthetic gene clusters

Bacteria of the phylum Acidobacteria are one of the most abundant bacterial across soil ecosystems, yet they are represented by comparatively few sequenced genomes, leaving gaps in our understanding of their metabolic diversity. Recently, genomes of Acidobacteria species with unusually large repertoires of biosynthetic gene clusters (BGCs) were reconstructed from grassland soil metagenomes, but the degree to which these species are widespread is still unknown. To investigate this, we augmented a dataset of publicly available Acidobacteria genomes with 46 metagenome-assembled genomes recovered from permanently saturated organic-rich soils of a vernal (spring) pool ecosystem in Northern California. We recovered high quality genomes for three novel species from Candidatus Angelobacter (a proposed subdivision 1 Acidobacterial genus), a genus that is genomically enriched in genes for specialized metabolite biosynthesis. Acidobacteria were particularly abundant in the vernal pool sediments, and a Ca. Angelobacter species was the most abundant bacterial species detected in some samples. We identified numerous diverse biosynthetic gene clusters in these genomes, and also in additional genomes from other publicly available soil metagenomes for other related Ca. Angelobacter species. Metabolic analysis indicates that Ca. Angelobacter likely are aerobes that ferment organic carbon, with potential to contribute to carbon compound turnover in soils. Using metatranscriptomics, we identified in situ expression of specialized metabolic traits for two species from this genus. In conclusion, we expand genomic sampling of the uncultivated Ca. Angelobacter, and show that they represent common and sometimes highly abundant members of dry and saturated soil communities, with a high degree of capacity for synthesis of diverse specialized metabolites.

microbiology↗

GIP: An open-source computational pipeline for mapping genomic instability from protists to cancer cells

Genome instability has been recognized as a key driver for microbial and cancer adaptation and thus plays a central role in many human pathologies. Even though genome instability encompasses different types of genomic alterations, most available genome analysis software are limited to just one kind mutation or analytical step. To overcome this limitation and better understand the role of genetic changes in enhancing pathogenicity we established GIP, a novel, powerful bioinformatic pipeline for comparative genome analysis. Here we show its application to whole genome sequencing datasets of Leishmania, Plasmodium, Candida, and cancer. Applying GIP on available data sets validated our pipeline and demonstrated the power of our analysis tool to drive biological discovery. Applied to Plasmodium vivax genomes, our pipeline allowed us to uncover the convergent amplification of erythrocyte binding proteins and to identify a nullisomic strain. Re-analyzing genomes of drug adapted Candida albicans strains revealed correlated copy number variations of functionally related genes, strongly supporting a mechanism of epistatic adaptation through interacting gene-dosage changes. Our results illustrate how GIP can be used for the identification of aneuploidy, gene copy number variations, changes in nucleic acid sequences, and chromosomal rearrangements. Altogether, GIP can shed light on the genetic bases of cell adaptation and drive disease biomarker discovery. One Sentence SummaryGIP - a novel pipeline for detecting, comparing and visualizing genome instability.

bioinformatics↗

Genomes of novel Myxococcota reveal severely curtailed machineries for predation and cellular differentiation

Cultured Myxococcota are predominantly aerobic soil inhabitants, characterized by their highly coordinated predation and cellular differentiation capacities. Little is currently known regarding yet-uncultured Myxococcota from anaerobic, non-soil habitats. We analyzed genomes representing one novel order (o__JAFGXQ01) and one novel family (f__JAFGIB01) in the Myxococcota from an anoxic freshwater spring in Oklahoma, USA. Compared to their soil counterparts, anaerobic Myxococcota possess smaller genomes, and a smaller number of genes encoding biosynthetic gene clusters (BGCs), peptidases, one- and two-component signal transduction systems, and transcriptional regulators. Detailed analysis of thirteen distinct pathways/processes crucial to predation and cellular differentiation revealed severely curtailed machineries, with the notable absence of homologs for key transcription factors (e.g. FruA and MrpC), outer membrane exchange receptor (TraA), and the majority of sporulation-specific and A-motility-specific genes. Further, machine-learning approaches based on a set of 634 genes informative of social lifestyle predicted a non-social behavior for Zodletone Myxococcota. Metabolically, Zodletone Myxococcota genomes lacked aerobic respiratory capacities, but encoded genes suggestive of fermentation, dissimilatory nitrite reduction, and dissimilatory sulfate-reduction (in f_JAFGIB01) for energy acquisition. We propose that predation and cellular differentiation represent a niche adaptation strategy that evolved circa 500 Mya in response to the rise of soil as a distinct habitat on earth. ImportanceThe Myxococcota is a phylogenetically coherent bacterial lineage that exhibits unique social traits. Cultured Myxococcoat are predominantly aerobic soil-dwelling microorganisms that are capable of predation and fruiting body formation. However, multiple yet-uncultured lineages within the Myxococcota has been encountered in a wide range of non-soil, predominantly anaerobic habitats; and the metabolic capabilities, physiological preferences, and capacity of social behavior of such lineages remains unclear. Here, we analyzed genomes recovered from a metagenomic analysis of an anoxic freshwater spring in Oklahoma, USA that represent novel, yet-uncultured, orders and families in the Myxococcota. The genomes appear to lack the characteristic hallmarks for social behavior encountered in Myxococcota genomes, and displayed a significantly smaller genome size and a smaller number of genes encoding biosynthetic gene clusters, peptidases, signal transduction systems, and transcriptional regulators. Such perceived lack of social capacity we confirmed through detailed comparative genomic analysis of thirteen pathways associated with Myxococcota social behavior, as well as the implementation of machine learning approaches to predict social behavior based on genome composition. Metabolically, these novel Myxococcota are predicted to be strict anaerobes, utilizing fermentation, nitrate rductio, and dissimilarity sulfate reduction for energy acquisition. Our result highlight the broad patterns of metabolic diversity within the yet-uncultured Myxococcota and suggest that the evolution of predation and fruiting body formation in the Myxococcoat has occurred in response to soil formation as a distinct habitat on earth.

microbiology↗

CoVizu: Rapid analysis and visualization of the global diversity of SARS-CoV-2 genomes

Phylogenetics has played a pivotal role in the genomic epidemiology of SARS-CoV-2, such as tracking the emergence and global spread of variants, and scientific communication. However, the rapid accumulation of genomic data from around the world -- with over two million genomes currently available in the GISAID database -- is testing the limits of standard phylogenetic methods. Here, we describe a new approach to rapidly analyze and visualize large numbers of SARS-CoV-2 genomes. Using Python, genomes are filtered for problematic sites, incomplete coverage, and excessive divergence from a strict molecular clock. All differences from the reference genome, including indels, are extracted using minimap2, and compactly stored as a set of features for each genome. For each Pango lineage (https://cov-lineages.org), we collapse genomes with identical features into variants, generate 100 bootstrap samples of the feature set union to generate weights, and compute the symmetric differences between the weighted feature sets for every pair of variants. The resulting distance matrices are used to generate neigihbor-joining trees in RapidNJ and converted into a majority-rule consensus tree for the lineage. Branches with support values below 50% or mean lengths below 0.5 differences are collapsed, and tip labels on affected branches are mapped to internal nodes as directly-sampled ancestral variants. Currently, we process about million genomes in approximately nine hours on 34 cores. The resulting trees are visualized using the JavaScript framework D3.js as beadplots, in which variants are represented by horizontal line segments, annotated with beads representing samples by collection date. Variants are linked by vertical edges to represent branches in the consensus tree. These visualizations are published at https://filogeneti.ca/CoVizu. All source code was released under an MIT license at https://github.com/PoonLab/covizu.

bioinformatics↗

Patterns of gene co-expression under water-deficit treatments and pan-genome occupancy in Brachypodium distachyon.

Natural populations are characterized by abundant genetic diversity driven by a range of different types of mutation. The tractability of sequencing complete genomes has allowed new insights into the variable composition of genomes, summarized as a species pan-genome. These analyses demonstrate that many genes are absent from the first reference genomes, whose analysis dominated the initial years of the genomic era. Our field now turns towards understanding the functional consequence of these highly variable genomes. Here, we analyzed weighted gene co-expression networks from leaf transcriptome data for drought response in the purple false brome Brachypodium distachyon and the differential expression of genes putatively involved in adaptation to this stressor. We specifically asked whether genes with variable "occupancy" in the pan-genome - genes which are either present in all studied genotypes or missing in some genotypes - show different distributions among co-expression modules. Co-expression analysis united genes expressed in drought-stressed plants into nine modules covering 72 hub genes (87 hub isoforms), and genes expressed under controlled water conditions into 13 modules, covering 190 hub genes (251 hub isoforms). We find that low occupancy pan-genes are under-represented among several modules, while other modules are over-enriched for low-occupancy pan-genes. We also provide new insight into the regulation of drought response in B. distachyon, specifically identifying one module with an apparent role in primary metabolism that is strongly responsive to drought. Our work shows the power of integrating pan-genomic analysis with transcriptomic data using factorial experiments to understand the functional genomics of environmental response.

plant biology↗

Trait-trait relationships and functional tradeoffs vary with genome size in prokaryotes

We report genomic traits that have been associated with the life history of prokaryotes and highlight conflicting findings concerning earlier observed trait correlations and tradeoffs. In order to address possible explanations for these contradictions we examined trait-trait variations of 11 genomic traits from ~ 18,000 sequenced genomes. The studied trait-trait variations suggested: (i) the predominance of two resistance and resilience-related orthogonal axes and (ii) at least in free living species with large effective population sizes whose evolution is little affected by genetic drift an overlap between a resilience axis and an axis of resource usage efficiency. These findings imply that resistance associated traits of prokaryotes are globally decoupled from resilience related traits and in the case of free-living communities also from resource use efficiencies associated traits. However, further inspection of pairwise scatterplots showed that resistance and resilience traits tended to be positively related for genomes up to roughly five million base pairs and negatively for larger genomes. This in turn may preclude a globally consistent assignment of prokaryote genomic traits to the competitor - stress-tolerator - ruderal (CSR) schema that sorts species depending on their location along disturbance and productivity gradients into three ecological strategies and may serve as an explanation for conflicting findings from earlier studies. All reviewed genomic traits featured significant phylogenetic signals and we propose that our trait table can be applied to extrapolate genomic traits from taxonomic marker genes. This will enable to empirically evaluate the assembly of these genomic traits in prokaryotic communities from different habitats and under different productivity and disturbance scenarios as predicted via the resistance-resilience framework formulated here.

ecology↗

The miR-430 locus with extreme promoter density is a transcription body organizer, which facilitates long range regulation in zygotic genome activation

In anamniote embryos the major wave of zygotic genome activation (ZGA) starts during the mid-blastula transition. This major wave of ZGA is facilitated by several mechanisms, including dilution of repressive maternal factors and accumulation of activating transcription factors during the fast cell division cycles preceding the mid-blastula transition. However, a set of genes escape global genome repression and are activated substantially earlier, during what is called, the minor wave of genome activation. While the mechanisms underlying the major wave of genome activation have been studied extensively, the minor wave of genome activation is little understood. In zebrafish the earliest expressed RNA polymerase II (Pol II) transcribed genes are activated in a pair of large transcription bodies depleted of chromatin, abundant in elongating Pol II and nascent RNAs (Hadzhiev et al., 2019; Hilbert et al., 2021). This transcription body includes the miR-430 gene cluster required for maternal mRNA clearance. Here we explored the genomic, chromatin organisation and cis-regulatory mechanisms of the minor wave of genome activation occurring in the transcription body. By long read genome sequencing we identified a remarkable cluster of miR-430 genes with over 300 promoters and spanning 0.6 Mb, which represent the highest promoter density of the genome. We demonstrate that the miR-430 gene cluster is required for the formation of the transcription body and acts as a transcription organiser for minor wave activation of a set of zinc finger genes scattered on the same chromosome arm, which share promoter features with the miR-430 cluster. These promoter features are shared among minor wave genes overall and include the TATA-box and sharp transcription start site profile. Single copy miR-430 promoter transgene reporter experiments indicate the importance of promoter-autonomous mechanisms regulating escape from global repression of the early embryo. These results together suggest that formation of the transcription body in the early embryo is the result of high promoter density coupled to a minor wave-specific core promoter code for transcribing key minor wave ZGA genes, which are required for the overhaul of the transcriptome during early embryonic development.

developmental biology↗

The OceanDNA MAG catalog contains over 50,000 prokaryotic genomes originated from various marine environments

Marine microorganisms are immensely diverse and play fundamental roles in global geochemical cycling. Recent metagenome-assembled genome studies, with special attention to large-scale projects such as Tara Oceans, have expanded the genomic repertoire of marine microorganisms. However, published marine metagenome data has not been fully explored yet. Here, we collected 2,057 marine metagenomes (>29 Tera bps of sequences) covering various marine environments and developed a new genome reconstruction pipeline. We reconstructed 52,325 qualified genomes composed of 8,466 prokaryotic species-level clusters spanning 59 phyla, including genomes from deep-sea deeper than 1,000 m (n=3,337), low-oxygen zones of <90 mol O2 per kg water (n=7,884), and polar regions (n=7,752). Novelty evaluation using a genome taxonomy database shows that 6,256 species (73.9%) are novel and include genomes of high taxonomic novelty such as new class candidates. These genomes collectively expanded the known phylogenetic diversity of marine prokaryotes by 34.2% and the species representatives cover 26.5 - 42.0% of prokaryote-enriched metagenomes. This genome resource, thoroughly leveraging accumulated metagenomic data, illuminates uncharacterized marine microbial dark matter lineages.

microbiology↗

A bacterial genome and culture collection of gut microbial in weanling piglet

The microbiota hosted in the pig gastrointestinal tract are important for productivity of livestock. However, the individual species and functional repertoires that make up the pig gut microbiome remain largely undefined. Here we comprehensively investigated the genomes and functions of the piglet gut microbiome using culture-based and metagenomics approaches. A collection included 266 cultured genomes and 482 metagenome-assembled genomes (MAGs) that were clustered to 428 species across 10 phyla was established. Among these clustered species, 333 genomes represent potential new species. Less matches between cultured genomes and MAGs revealed a substantial bias for the acquisition of reference genomes by the two strategies. Glycoside hydrolases was the dominant category of carbohydrate-active enzymes. 445 secondary metabolite biosynthetic genes were predicted from 292 genomes with bacteriocin being the most. Pan genome analysis of Limosilactobacillus reuteri uncover the biosynthesis of reuterin was strain-specific and the production was experimentally determined. These genomic resources will enable a comprehensive characterization of the microbiome composition and function of pig gut.

microbiology↗

Conserved genomic landscapes of differentiation across Populus speciation continuum

Speciation, the continuous process by which new species form, is often investigated by looking at the variation of nucleotide diversity and differentiation across the genome (hereafter genomic landscapes). A key challenge lies in how to determine the main evolutionary forces at play shaping these patterns. One promising strategy, albeit little used to date, is to comparatively investigate these genomic landscapes as a progression through time by using a series of species pairs along a divergence gradient. Here, we resequenced 201 whole-genomes from eight closely related Populus species, with pairs of species at different stages along the divergence gradient to learn more about speciation processes. Using population structure and ancestry analyses, we document extensive introgression between some species pairs, especially those with parapatric distributions. We further investigate genomic landscapes, focusing on within-species (nucleotide diversity and recombination rate) and among-species (relative and absolute divergence) summary statistics of diversity and divergence. We observe highly conserved patterns of genomic divergence across species pairs. Independent of the stage across the divergence gradient, we find support for signatures of linked selection (i.e., the interaction between natural selection and genetic linkage) in shaping these genomic landscapes, along with gene flow and standing genetic variation. We highlight the importance of investigating genomic patterns on multiple species across a divergence gradient and discuss prospects to better understand the evolutionary forces shaping the genomic landscapes of diversity and differentiation.

evolutionary biology↗

Short '1.2x genome' infectious clone initiates deltavirus replication in Boa constrictor cells

Human hepatitis D virus (HDV), discovered in 1977, represented the sole known deltavirus for decades. The dependence on hepatitis B virus (HBV) co-infection and its glycoproteins for infectious particle formation led to the assumption that deltaviruses are human-only pathogens. However, since 2018, several reports have described identification of HDV-like agents from various hosts but without co-infecting hepadnaviruses. Indeed, we demonstrated that Swiss snake colony virus 1 (SwSCV-1) uses arenaviruses as the helper for infectious particle formation, thus shaking the dogmatic alliance with hepadnaviruses for completing deltavirus life cycle. In vitro systems enabling helper virus-independent replication are key for studying the newly discovered deltaviruses. Others and we have successfully used constructs containing multimers of the deltavirus genome for the replication of various deltaviruses via transfection in cell culture. Here, we report the establishment of deltavirus infectious clones with 1.2x genome inserts bearing two copies of the genomic and antigenomic ribozymes. We used SwSCV-1 as the model to compare the ability of the previously reported "2x genome" and the "1.2x genome" plasmid constructs/infectious clones to initiate replication in cell culture. Using immunofluorescence, qRT-PCR, immuno- and northern blotting, we found the 2x and 1.2x genome clones to similarly initiate deltavirus replication in vitro and both induced a persistent infection of snake cells. We hypothesize that duplicating the ribozymes facilitates the cleavage of genome multimers into unit-length pieces during the initial round of replication. The 1.2x genome constructs enable easier introduction of modifications required for studying deltavirus replication and cellular interactions. IMPORTANCEHepatitis D virus (HDV) is a satellite virus infecting humans with strict association to hepatitis B virus (HBV) co-infection because HBV glycoproteins can mediate infectious HDV particle formation. For decades, HDV was the sole representative of deltaviruses, which had led to hypotheses suggesting that it evolved in humans, the only known natural host. Recent sequencing studies have led to the discovery of HDV-like sequences across a wide range of species, representing a paradigm shift in deltavirus evolution. Molecular biology tools such as infectious clones, which enable initiation of deltavirus infection without helper virus, are key to demonstrate that the recently found deltaviruses are capable of independent replication. Such tools will enable identification of the potential helper viruses. Here, we report a 1.2x genome copy strategy for designing plasmid-based infectious clones to study deltaviruses and to demonstrate that plasmid delivery into cultured snake cells sufficiently initiates replication of different deltaviruses.

molecular biology↗

Deep mining of the Sequence Read Archive reveals bipartite coronavirus genomes and inter-family Spike glycoprotein recombination

Genetic variation in RNA viruses is generated by point mutation and recombination as well as reassortment in the case of viruses with segmented genomes. While point mutation concerns only few sites per genome copy, recombination and reassortment can affect large genome regions, possibly facilitating the sudden emergence of novel traits. The contribution of recombination and reassortment to genomic plasticity and their rates remain poorly understood and might be underappreciated because of the lack of a comprehensive description of the virosphere. Here we employed a computational approach that directly queries primary sequencing data in a highly parallelized way and involves a targeted viral genome assembly strategy. By screening more than 213,000 data sets from the Sequence Read Archive repository and using two metrics that quantitatively assess assembly quality we discovered 25 novel nidoviruses from a wide range of vertebrate hosts. These include eight fish coronaviruses with bipartite genomes, a giant 36.1 kilobase coronavirus genome with a duplicated Spike glycoprotein (S) gene, and 16 additional so far undescribed vertebrate nidoviruses. Some of these novel virus genomes encode protein domains that have not been described for nidoviruses. We provide evidence for a possible inter-family homologous recombination event involving S between ancestral bipartite coronaviruses and unsegmented tobaniviruses and report a case example of an individual fish simultaneously infected with members from both virus families. Our results shed light on the evolution and genomic plasticity of coronaviruses and identify recombinants with a possibly improved ability to cross species barriers, which might elevate their pandemic potential.

microbiology↗

A new lineage of non-photosynthetic green algae with extreme organellar genomes

BackgroundThe plastid genomes of the green algal order Chlamydomonadales tend to expand their non-coding regions, but this phenomenon is poorly understood. Here we shed new light on organellar genome evolution in Chlamydomonadales by studying a previously unknown non-photosynthetic lineage. We established cultures of two new Polytoma-like flagellates, defined their basic characteristics and phylogenetic position, and obtained complete organellar genome sequences and a transcriptome assembly for one of them. ResultsWe discovered a novel deeply diverged chlamydomonadalean lineage that has no close photosynthetic relatives and represents an independent case of photosynthesis loss. To accommodate these organisms we establish the new genus Leontynka, with two species (L. pallida and L. elongata) distinguishable through both their morphological and molecular characteristics. Notable features of the colourless plastid of L. pallida deduced from the plastid genome (plastome) sequence and transcriptome assembly include the retention of ATP synthase, thylakoid-associated proteins, the carotenoid biosynthesis pathway, and a plastoquinone-based electron transport chain, the latter two modules having an obvious functional link to the eyespot present in Leontynka. Most strikingly, the ~362 kbp plastome of L. pallida is by far the largest among the non-photosynthetic eukaryotes investigated to date due to an extreme proliferation of sequence repeats. These repeats are also present in coding sequences, with one repeat type found in the exons of 11 out of 34 protein-coding genes, with up to 36 copies per gene, thus affecting the encoded proteins. The mitochondrial genome of L. pallida is likewise exceptionally large, with its >104 kbp surpassed only by the mitogenome of Haematococcus lacustris among all members of Chlamydomonadales hitherto studied. It is also bloated with repeats, though entirely different from those in the L. pallida plastome, which contrasts with the situation in H. lacustris where both the organellar genomes have accumulated related repeats. Furthermore, the L. pallida mitogenome exhibits an extremely high GC content in both coding and non-coding regions and, strikingly, a high number of predicted G-quadruplexes. ConclusionsWith its unprecedented combination of plastid and mitochondrial genome characteristics, Leontynka pushes the frontiers of organellar genome diversity and is an interesting model for studying organellar genome evolution.

evolutionary biology↗

Rapid genomic evolution in Brassica rapa with bumblebee selection in experimental evolution

BackgroundInsect pollinators shape rapid phenotypic evolution of traits related to floral attractiveness and plant reproductive success. However, the underlying genomic changes remain largely unknown despite their importance in predicting adaptive responses to natural or to artificial selection. Based on a nine-generation experimental evolution study with fast cycling Brassica rapa plants adapting to bumblebees, we investigate the genomic evolution associated with the previously observed parallel phenotypic evolution. In this current evolve and resequencing (E&R) study, we conduct a genomic scan of the allele frequency changes along the genome in bumblebee-pollinated and hand-pollinated plants and perform a genomic principal component analysis (PCA). ResultsWe highlight rapid genomic evolution associated with the observed phenotypic evolution mediated by bumblebees. Controlling for genetic drift, we observe significant changes in allelic frequencies at multiple loci. However, this pattern differs according to the replicate of bumblebee-pollinated plants, suggesting putative non-parallel genomic evolution. Finally, our study underlines an increase in genomic differentiation implying the putative involvement of multiple loci in short-term pollinator adaptation. ConclusionsOverall, our study enhances our understanding of the complex interactions between pollinator and plants, providing a steppingstone towards unravelling the genetic basis of plant genomic adaptation to biotic factors in the environment.

evolutionary biology↗

binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets

The reconstruction of genomes is a critical step in genome-resolved metagenomics and for multi-omic data integration from microbial communities. Here, we present binny, a binning tool that produces complete and pure metagenome-assembled genomes (MAG) from both contiguous and highly fragmented genomes. Based on established metrics, binny outperforms or is highly competitive with commonly-used and state- of-the-art binning methods and finds unique genomes that could not be detected by other methods. binny uses k-mer-composition and coverage by metagenomic reads for iterative, non-linear dimension reduction of genomic signatures, as well as subsequent automated contig clustering with cluster assessment using lineage-specific marker gene sets. When compared to seven widely used binning algorithms, binny provides substantial amounts of uniquely identified MAGs and almost always recovers the most near-complete (>95% pure, >90% complete) and high-quality (>90% pure, >70% complete) genomes from simulated data sets from the Critical Assessment of Metagenome Interpretation (CAMI) initiative, as well as substantially more high-quality draft genomes, as defined by the Minimum Information about a Metagenome-Assembled Genome (MIMAG) standard, from a real-world benchmark comprised of metagenomes from various environments than any other tested method.

bioinformatics↗