bioRxiv ScienceSearch

Biology subjects

Clavijo, B. J.

Publications and source records attributed to Clavijo, B. J..

3 recordsLinked to original sources

Skip-mers: increasing entropy and sensitivity to detect conserved genic regions with simple cyclic q-grams

Bioinformatic analyses and tools make extensive use of k-mers (fixed contiguous strings of k nucleotides) as an informational unit. K-mer analyses are both useful and fast, but are strongly affected by single-nucleotide polymorphisms or sequencing errors, effectively hindering direct-analyses of whole regions and decreasing their usability between evolutionary distant samples.\n\nWe introduce a concept of skip-mers, a cyclic pattern of used-and-skipped positions of k nucleotides spanning a region of size S [≥] k, and show how analyses are improved compared to using k-mers. The entropy of skip-mers increases with the larger span, capturing information from more distant positions and increasing the specificity, and uniqueness, of larger span skip-mers within a genome. In addition, skip-mers constructed in cycles of 1 or 2 nucleotides in every 3 (or a multiple of 3) lead to increased sensitivity in the coding regions of genes, by grouping together the more conserved nucleotides of the protein-coding regions.\n\nWe implemented a set of tools to count and intersect skip-mers between different datasets. We used these tools to show how skip-mers have advantages over k-mers in terms of entropy and increased sensitivity to detect conserved coding sequence, allowing better identification of genic matches between evolutionarily distant species. We also highlight potential applications to problems such as whole-genome alignment and multi-genome evolutionary analyses.\n\nSoftware availabilitythe skm-tools implementing the methods described in this manuscript are available under MIT license at http://github.com/bioinfologics/skm-tools/

bioinformatics

The ash dieback invasion of Europe was founded by two individuals from a native population with huge adaptive potential

Accelerating international trade and climate change make pathogen spread an increasing concern. Hymenoscyphus fraxineus, the causal agent of ash dieback is one such pathogen, moving across continents and hosts from Asian to European ash. Most European common ash (Fraxinus excelsior) trees are highly susceptible to H. fraxineus although a small minority (~5%) evidently have partial resistance to dieback. We have assembled and annotated a draft of the H. fraxineus genome which approaches chromosome scale. Pathogen genetic diversity across Europe, and in Japan, reveals a tight bottleneck into Europe, though a signal of adaptive diversity remains in key host interaction genes (effectors). We find that the European population was founded by two divergent haploid individuals. Divergence between these haplotypes represents the 'shadow' of a large source population and subsequent introduction would greatly increase adaptive potential and the pathogen's threat. Thus, EU wide biological security measures remain an important part of the strategy to manage this disease.

genomics

An improved assembly and annotation of the allohexaploid wheat genome identifies complete families of agronomic genes and provides genomic evidence for chromosomal translocations.

Advances in genome sequencing and assembly technologies are generating many high quality genome sequences, but assemblies of large, repeat-rich polyploid genomes, such as that of bread wheat, remain fragmented and incomplete. We have generated a new wheat whole-genome shotgun sequence assembly using a combination of optimised data types and an assembly algorithm designed to deal with large and complex genomes. The new assembly represents more than 78% of the genome with a scaffold N50 of 88.8kbp that has a high fidelity to the input data. Our new annotation combines strand-specific Illumina RNAseq and PacBio full-length cDNAs to identify 104,091 high confidence protein-coding genes and 10,156 non-coding RNA genes. We confirmed three known and identified one novel genome rearrangements. Our approach enables the rapid and scalable assembly of wheat genomes, the identification of structural variants, and the definition of complete gene models, all powerful resources for trait analysis and breeding of this key global crop. [Supplemental material is available for this article.]

genomics