bioRxiv ScienceSearch

Biology subjects

Heitkam, T.

Publications and source records attributed to Heitkam, T..

6 recordsLinked to original sources

Complete pan-plastome sequences enable high resolution phylogenetic classification of sugar beet and closely related crop wild relatives

BackgroundAs the major source of sugar in moderate climates, sugar-producing beets (Beta vulgaris subsp. vulgaris) have a high economic value. However, the low genetic diversity within cultivated beets requires introduction of new traits, for example to increase their tolerance and resistance attributes - traits that often reside in the crop wild relatives. For this, genetic information of wild beet relatives and their phylogenetic placements to each other are crucial. To answer this need, we sequenced and assembled the complete plastome sequences from a broad species spectrum across the beet genera Beta and Patellifolia, both embedded in the Betoideae (order Caryophyllales). This pan-plastome dataset was then used to determine the wild beet phylogeny in high-resolution. ResultsWe sequenced the plastomes of 18 closely related accessions representing 11 species of the Betoideae subfamily and provided high-quality plastome assemblies which represent an important resource for further studies of beet wild relatives and the diverse plant order Caryophyllales. Their assembly sizes range from 149,723 bp (Beta vulgaris subsp. vulgaris) to 152,816 bp (Beta nana), with most variability in the intergenic sequences. Combining plastome-derived phylogenies with read-based treatments based on mitochondrial information, we were able to suggest a unified and highly confident phylogenetic placement of the investigated Betoideae species. Our results show that the genus Beta can be divided into the two clearly separated sections Beta and Corollinae. Our analysis confirms the affiliation of B. nana with the other Corollinae species, and we argue against a separate placement in the Nanae section. Within the Patellifolia genus, the two diploid species Patellifolia procumbens and Patellifolia webbiana are, regarding the plastome sequences, genetically more similar to each other than to the tetraploid Patellifolia patellaris. Nevertheless, all three Patellifolia species are clearly separated. ConclusionIn conclusion, our wild beet plastome assemblies represent a new resource to understand the molecular base of the beet germplasm. Despite large differences on the phenotypic level, our pan-plastome dataset is highly conserved. For the first time in beets, our whole plastome sequences overcome the low sequence variation in individual genes and provide the molecular backbone for highly resolved beet phylogenomics. Hence, our plastome sequencing strategy can also guide genomic approaches to unravel other closely related taxa.

genomics

Genome-wide analysis of long terminal repeat retrotransposons from the cranberry Vaccinium macrocarpon

BACKGROUNDLong terminal repeat (LTR) retrotransposons are widespread in plant genomes and play a large role in the generation of genomic variation. Despite this, their identification and characterization remains challenging, especially for non-model genomes. Hence, LTR retrotransposons remain undercharacterized in Vaccinium genomes, although they may be beneficial for current berry breeding efforts. OBJECTIVEExemplarily focusing on the genome of American cranberry (Vaccinium macrocarpon Aiton), we aim to generate an overview of the LTR retrotransposon landscape, highlighting the abundance, transcriptional activity, sequence, and structure of the major retrotransposon lineages. METHODSGraph-based clustering of whole genome shotgun Illumina reads was performed to identify the most abundant LTR retrotransposons and to reconstruct representative in silico full-length elements. To generate insights into the LTR retrotransposon diversity in V. macrocarpon, we also queried the genome assembly for presence of reverse transcriptases (RTs), the key domain of LTR retrotransposons. Using transcriptomic data, transcriptional activity of retrotransposons corresponding to the consensuses was analyzed. RESULTSWe provide an in-depth characterization of the LTR retrotransposon landscape in the V. macrocarpon genome. Based on 475 RTs harvested from the genome assembly, we detect a high retrotransposon variety, with all major lineages present. To better understand their structural hallmarks, we reconstructed 26 Ty1-copia and 28 Ty3-gypsy in silico consensuses that capture the detected diversity. Accordingly, we frequently identify association with tandemly repeated motifs, extra open reading frames, and specialized, lineage-typical domains. Based on the overall high genomic abundance and transcriptional activity, we suggest that retrotransposons of the Ale and Athila lineages are most promising to monitor retrotransposon-derived polymorphisms across accessions. CONCLUSIONSWe conclude that LTR retrotransposons are major components of the V. macrocarpon genome. The representative consensuses provide an entry point for further Vaccinium genome analyses and may be applied to derive molecular markers for enhancing cranberry selection and breeding.

plant biology

ECCsplorer: a pipeline to detect extrachromosomal circular DNA (eccDNA) from next-generation sequencing data

MotivationExtrachromosomal circular DNAs (eccDNAs) are ring-like DNA structures physically separated from the chromosomes with 100 bp to several megabasepairs in size. Apart from carrying tandemly repeated DNA, eccDNAs may also harbor extra copies of genes or recently activated transposable elements. As eccDNAs occur in all eukaryotes investigated so far and likely play roles in stress, cancer, and aging, they have been prime targets in recent research - with their investigation limited by the scarcity of computational tools. ResultsHere, we present the ECCsplorer, a bioinformatics pipeline to detect eccDNAs in any kind of organism or tissue using next-generation sequencing techniques. Following Illumina-sequencing of amplified circular DNA (circSeq), the ECCsplorer enables an easy and automated discovery of eccDNA candidates. The data analysis encompasses two major procedures: First, read mapping to the reference genome allows the detection of informative read distributions including high coverage, discordant mapping, and split reads. Second, reference-free comparison of read clusters from amplified eccDNA against control sample data reveals specifically enriched DNA circles. Both software parts can be run separately or jointly, depending on the individual aim or data availability. To illustrate the wide applicability of our approach, we analyzed semiartificial and published circSeq data from the model organisms H. sapiens and A. thaliana, and generated circSeq reads from the non-model crop B. vulgaris. We clearly identified eccDNA candidates from all datasets, with and without reference genomes. The ECCsplorer pipeline specifically detected mitochondrial mini-circles and retrotransposon activation, showcasing the ECCsplorers sensitivity and specificity. The derived eccDNA targets are valuable for a wide range of downstream investigations - from analysis of cancer-related eccDNAs over organelle genomics to identification of active transposable elements. Availability and implementationThe ECCsplorer pipeline is available on GitHub at https://github.com/crimBubble/ECCsplorer under the GNU license. ContactTony Heitkam (tony.heitkam@tu-dresden.de) Supplementary informationSupplementary data are available online.

bioinformatics

Comparative repeat profiling of two closely related conifers (Larix decidua and Larix kaempferi) reveals high genome similarity with only one fast-evolving satellite DNA

In eukaryotic genomes, cycles of repeat expansion and removal lead to large-scale genomic changes and propel organisms forward in evolution. However, in conifers, active repeat removal is thought to be limited, leading to expansions of their genomes, mostly exceeding 10 gigabasepairs. As a result, conifer genomes are largely littered with fragmented and decayed repeats. Here, we aim to investigate how the repeat landscapes of two related conifers have diverged, given the conifers accumulative genome evolution mode. For this, we applied low coverage sequencing and read clustering to the genomes of European and Japanese larch, Larix decidua (Lamb.) Carriere and Larix kaempferi (Mill.), that arose from a common ancestor, but are now geographically isolated. We found that both Larix species harbored largely similar repeat landscapes, especially regarding the transposable element content. To pin down possible genomic changes, we focused on the repeat class with the fastest sequence turnover: satellite DNAs (satDNAs). Using comparative bioinformatics, Southern, and fluorescent in situ hybridization, we reveal the satDNAs organizational patterns, their abundances, and chromosomal locations. Four out of the five identified satDNAs are widespread in the Larix genus, with two even present in the more distantly related Pseudotsuga and Abies genera. Unexpectedly, the EulaSat3 family was restricted to L. decidua and absent from L. kaempferi, indicating its evolutionarily young age. Taken together, our results exemplify how the accumulative genome evolution of conifers may limit the overall divergence of repeats after speciation, producing only few repeat-induced genomic novelties.

plant biology

Broken, silent, and in hiding: Tamed endogenous pararetroviruses escape elimination from the genome of sugar beet (Beta vulgaris)

Background and AimsEndogenous pararetroviruses (EPRVs) are widespread components of plant genomes that originated from episomal DNA viruses of the Caulimoviridae family. Due to fragmentation and rearrangements, most EPRVs have lost their ability to replicate through reverse transcription and to initiate viral infection. Similar to the closely related retrotransposons, extant EPRVs were retained and often amplified in plant genomes for several million years. Here, we characterize the complete genomic EPRV fraction of the crop sugar beet (Beta vulgaris, Amaranthaceae) to understand how they shaped the beet genome and to suggest explanations for their absent virulence. MethodsUsing next- and third-generation sequencing data and the genome assembly, we reconstructed full-length in silico representatives for the three host-specific EPRV families (beetEPRVs) in the B. vulgaris genome. Focusing on the canonical family beetEPRV3, we investigated its chromosomal localization, abundance, and distribution by fluorescent in situ and Southern hybridization. Key ResultsBeetEPRVs range between 7.5 and 10.7 kb (0.3 % of the B. vulgaris genome) and are heterogeneous in structure and sequence. Although all three beetEPRV families were assigned to the florendoviruses, they showed variably arranged protein-coding domains, different degrees of fragmentation, and preferences for diverse sequence contexts. We observed small RNAs that target beetEPRVs in a family-specific manner, indicating stringent epigenetic suppression. We localized beetEPRV3 on all 18 sugar beet chromosomes, occurring preferentially in clusters and associated with heterochromatic, centromeric and intercalary satellite DNAs. BeetEPRV3 variants also exist in the genomes of related wild species, indicating an initial beetEPRV3 integration 13.4 to 7.2 million years ago. ConclusionsOur study in beet illustrates the variability of EPRV structure and sequence in a single host genome. Evidence of sequence fragmentation and epigenetic silencing imply possible plant strategies to cope with long-term persistence of EPRVs, including amplification, fixation in the heterochromatin, and containment of EPRV virulence.

plant biology

Satellite DNA landscapes after allotetraploidisation of quinoa (Chenopodium quinoa) reveal unique A and B subgenomes

If two related plant species hybridise, their genomes are combined within a single nucleus, thereby forming an allotetraploid. How the emerging plant balances two co-evolved genomes is still a matter of ongoing research. Here, we focus on satellite DNA (satDNA), the fastest turn-over sequence class in eukaryotes, aiming to trace its emergence, amplification and loss during plant speciation and allopolyploidisation. As a model, we used Chenopodium quinoa Willd. (quinoa), an allopolyploid crop with 2n=4x=36 chromosomes. Quinoa originated by hybridisation of an unknown female American Chenopodium diploid (AA genome) with an unknown male Old World diploid species (BB genome), dating back 3.3 to 6.3 million years. Applying short read clustering to quinoa (AABB), C. pallidicaule (AA), and C. suecicum (BB) whole genome shotgun sequences, we classified their repetitive fractions, and identified and characterised seven satDNA families, together with the 5S rDNA model repeat. We show unequal satDNA amplification (two families) and exclusive occurrence (four families) in the AA and BB diploids by read mapping as well as Southern, genomic and fluorescent in situ hybridisation. As C. pallidicaule harbours a unique satDNA profile, we are able to exclude it as quinoas parental species. Using quinoa long reads and scaffolds, we detected only limited evidence of interlocus homogenisation of satDNA after allopolyploidisation, but were able to exclude dispersal of 5S rRNA genes between subgenomes. Our results exemplify the complex route of tandem repeat evolution through Chenopodium speciation and allopolyploidisation, and may provide sequence targets for the identification of quinoas progenitors.

plant biology