bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Cost-effective long-read assembly of a hybrid Formica aquilonia x Formica polyctena wood ant genome from a single haploid individual

Formica red wood ants are a keystone species of boreal forest ecosystems and an emerging model system in the study of speciation and hybridization. Here we performed a standard DNA extraction from a single, field-collected Formica aquilonia x Formica polyctena haploid male and assembled its genome using [~]60x of PacBio long reads. After polishing and contaminant removal, the final assembly was 272 Mb (4,687 contigs, N50 = 1.16 Mb). Our reference genome contains 98.5% of the core Hymenoptera BUSCOs and was scaffolded using the pseudo-chromosomal assembly of a related species, F. selysi (28 scaffolds, N50 = 8.49 Mb). Around one third of the genome consists of repeats, and 17,426 gene models were annotated using both protein and RNAseq data (97.4% BUSCO completeness). This resource is of comparable quality to the few other single individual insect genomes assembled to date and paves the way to genomic studies of admixture in natural populations and comparative genomic approaches in Formica wood ants.

genomics

Genome features of common vetch (Vicia sativa) in natural habitats

Wild plants are often tolerant to biotic and abiotic stresses in their natural environments, whereas domesticated plants such as crops frequently lack such resilience. This difference is thought to be due to the high levels of genome heterozygosity in wild plant populations and the low levels of heterozygosity in domesticated crop species. In this study, common vetch (Vicia sativa) was used as a model to examine this hypothesis. The common vetch genome (2n = 14) was estimated as 1.8 Gb in size. Genome sequencing produced a reference assembly that spanned 1.5 Gb, from which 31,146 genes were predicted. Using this sequence as a reference, 24,118 single nucleotide polymorphisms were discovered in 1,243 plants from 12 natural common vetch populations in Japan. Common vetch genomes exhibited high heterozygosity at the population level, with lower levels of heterozygosity observed at specific genome regions. Such patterns of heterozygosity are thought to be essential for adaptation to different environments. These findings suggest that high heterozygosity at the population level would be required for wild plants to survive under natural conditions while allowing important gene loci to be fixed to adapt the conditions. The resources generated in this study will provide insights into de novo domestication of wild plants and agricultural enhancement. HighlightSequence analysis of the common vetch (Vicia sativa) genome and SNP genotyping across natural populations revealed nucleotide diversity levels associated with native population environments.

genomics

Accurate analysis of short read sequencing in complex genomes: A case study using QTL-seq to target blanchability in peanut (Arachis hypogaea)

Next Generation sequencing was a step change for molecular genetics and genomics. Illumina sequencing in particular still provides substantial value to animal and plant genomics. A simple yet powerful technique, referred to as QTL sequencing (QTL-seq) is susceptible to high levels of noise due to ambiguity of alignment of short reads in complex regions of the genome. This noise is particularly high when working with polyploid and/or outcrossing crop species, which impairs the efficacy of QTL-seq in identifying functional variation. By filtering loci based on the optimal alignment of short reads, we have developed a pipeline, named Khufu, that substantially improves the accuracy of QTL-seq analysis in complex genomes, allowing de novo variant discovery directly from bulk sequence. We first demonstrate the pipeline by identifying and validating loci contributing to blanching percentage in peanut using lines from multiple related populations. Using other published datasets in peanut, Brassica rapa, Hordeum volgare, Lactua satvia, and Felis catus, we demonstrate that Khufu produces more accurate results straight from bulk sequence. Khufu works across species, genome ploidy level, and data types. In cases where identified QTL were fine mapped, the fine mapped region corresponds to the top of the peak identified by Khufu. The accuracy of Khufu allows the analysis of population sequencing at very low coverage (<3x), greatly decreasing the amount of sequence needed to genotype even the most complex genomes.

genomics

Transcript- and annotation-guided genome assembly of the European starling

The European starling, Sturnus vulgaris, is an ecologically significant, globally invasive avian species that is also suffering from a major decline in its native range. Here, we present the genome assembly and long-read transcriptome of an Australian-sourced European starling (S. vulgaris vAU), and a second North American genome (S. vulgaris vNA), as complementary reference genomes for population genetic and evolutionary characterisation. S. vulgaris vAU combined 10x Genomics linked-reads, low-coverage Nanopore sequencing, and PacBio Iso-Seq full-length transcript scaffolding to generate a 1050 Mb assembly on 1,628 scaffolds (72.5 Mb scaffold N50). Species-specific transcript mapping and gene annotation revealed high structural and functional completeness (94.6% BUSCO completeness). Further scaffolding against the high-quality zebra finch (Taeniopygia guttata) genome assigned 98.6% of the assembly to 32 putative nuclear chromosome scaffolds. Rapid, recent advances in sequencing technologies and bioinformatics software have highlighted the need for evidence-based assessment of assembly decisions on a case-by-case basis. Using S. vulgaris vAU, we demonstrate how the multifunctional use of PacBio Iso-Seq transcript data and complementary homology-based annotation of sequential assembly steps (assessed using a new tool, SAAGA) can be used to assess, inform, and validate assembly workflow decisions. We also highlight some counter-intuitive behaviour in traditional BUSCO metrics, and present BO_SCPLOWUSCOMPC_SCPLOW, a complementary tool for assembly comparison designed to be robust to differences in assembly size and base-calling quality. Finally, we present a second starling assembly, S. vulgaris vNA, to facilitate comparative analysis and global genomic research on this ecologically important species.

genomics

Chromosome-level genome assembly of Acanthopagrus latus using PacBio and Hi-C technologies

The yellowfin seabream, Acanthopagrus latus, is widely distributed throughout the Indo-West Pacific. This fish is an ideal model species in which to study the mechanism of sex reversal since it exhibits a specific feature: sequential hermaphrodite. Here, we report a chromosome-scale assembly of the A. latus based on PacBio and Hi-C data. 22,485 protein-coding genes were annotated in whole genome level using transcriptome data. Taken together, this highly accurate, chromosome-level reference genome can provide a valuable resource to elucidate the mechanism of sex reversal for A. latus. Background & SummaryEvolution of sex, especially the evolution of different sexual systems, is a fascinating subject in evolutionary biology. The Sparidae, commonly known as seabreams or progies, is a family of fishes of the order Perciformes. And this family consist about 150 species in the world, which are mainly coastal fish1. Previous researchers mentioned that Sparidae is an ideal taxon to study the evolution history and adaptive significance of sexual systems, particularly for both types of sequential hermaphroditism, given that this group contains many protogyny, protandry and genochorist species2. The yellowfin seabream, Acanthopagrus latus is a protandry species which belongs to the Sparidae family. It is widely distributed in Indo-West Pacific area3. It has a great relevance for marine aquaculture and its biology is well focused on reproductive physiology and nutrition4. Interestingly, A. latus has a special gender feature is that it belongs to protandrous sexual system (initially as male and change later to female)5. Most of the past studies of A. latus mainly focused on the reproductive biology, population structure, aquaculture and taxonomy3,4,6-8. Although some sex reversal related genes were found in A. latus, the lack of genomic resources still limit us to elucidate the mechanism of sex reversal for this species.9,10. In addition, this lack was also limited the studies of evolution of sexual systems for Sparidae. In this study, long-read (PacBio SMRT) sequencing and Hi-C sequencing technologies were applied to construct a high quality reference genome for yellowfin seabream. This high-quality genome can provide a valuable resource to elucidate the mechanism of sex reversal for A. latus. Furthermore, this genome can also facilitate the studies of evolution of sexual systems for Sparidae.

genomics

Widespread false gene gains caused by duplication errors in genome assemblies

False duplications in genome assemblies lead to false biological conclusions. We quantified false duplications in previous genome assemblies and their new counterparts of the same species (platypus, zebra finch, Annas hummingbird) generated by the Vertebrate Genomes Project (VGP). Whole genome alignments revealed that 4 to 16% of the sequences were falsely duplicated in the previous assemblies, impacting hundreds to thousands of genes. These led to overestimated gene family expansions. The main source of the false duplications was heterotype duplications, where the haplotype sequences were more divergent than other parts of the genome leading the assembly algorithms to classify them as separate genes or genomic regions. A minor source was sequencing errors. Although present in a smaller proportion, we observed false duplications remaining in the VGP assemblies that can be identified and purged. This study highlights the need for more advanced assembly methods that better separates haplotypes and sequence errors, and the need for cautious analyses on gene gains.

genomics

Genome assembly of the popular Korean soybean cultivar Hwangkeum

Massive resequencing efforts have been undertaken to catalog allelic variants in major crop species including soybean, but the scope of the information for genetic variation often depends on short sequence reads mapped to the extant reference genome. Additional de novo assembled genome sequences provide a unique opportunity to explore a dispensable genome fraction in the pan-genome of a species. Here, we report the de novo assembly and annotation of Hwangkeum, a popular soybean cultivar in Korea. The assembly was constructed using PromethION nanopore sequencing data and two genetic maps, and was then error-corrected using Illumina short-reads and PacBio SMRT reads. The 933.12 Mb assembly was annotated 79,870 transcripts for 58,550 genes using RNA-Seq data and the public soybean annotation set. Comparison of the Hwangkeum assembly with the Williams 82 soybean reference genome sequence revealed 1.8 million single-nucleotide polymorphisms, 0.5 million indels, and 25 thousand putative structural variants. However, there was no natural megabase-scale chromosomal rearrangement. Incidentally, by adding two novel groups, we found that soybean contains four clearly separated groups of centromeric satellite repeats. Analyses of satellite repeats and gene content suggested that the Hwangkeum assembly is a high-quality assembly. This was further supported by comparison of the marker arrangement of anthocyanin biosynthesis genes and of gene arrangement at the Rsv3 locus. Therefore, the results indicate that the de novo assembly of Hwangkeum is a valuable additional reference genome resource for characterizing traits for the improvement of this important crop species.

genomics

Genome annotation with long RNA reads reveals new patterns of gene expression in an ant brain

Functional genomic analyses rely on high-quality genome assemblies and annotations. Highly contiguous genome assemblies have become available for a variety of species, but accurate and complete annotation of gene models, inclusive of alternative splice isoforms and transcription start and termination sites remains difficult with traditional approaches. Here, we utilized full-length isoform sequencing (Iso-Seq), a long-read RNA sequencing technology, to obtain a comprehensive annotation of the transcriptome of the ant Harpegnathos saltator. The improved genome annotations include additional splice isoforms and extended 3 untranslated regions for more than 4,000 genes. Reanalysis of RNA-seq experiments using these annotations revealed several genes with caste-specific differential expression and tissue-or caste-specific splicing patterns that were missed in previous analyses. The extended 3 untranslated regions afforded great improvements in the analysis of existing single-cell RNA-seq data, resulting in the recovery of the transcriptomes of 18% more cells. The deeper single-cell transcriptomes obtained with these new annotations allowed us to identify additional markers for several cell types in the ant brain, as well as genes differentially expressed across castes in specific cell types. Our results demonstrate that Iso-Seq is an efficient and effective approach to improve genome annotations and maximize the amount of information that can be obtained from existing and future genomic datasets in Harpegnathos and other organisms.

genomics

The genome of New Zealand trevally (Carangidae: Pseudocaranx georgianus) uncovers a XY sex determination locus

BackgroundThe genetic control of sex determinism in teleost species is poorly understood. This is partly because of the diversity of sex determining mechanisms in this large group, including constitutive genes linked to sex chromosomes, polygenic constitutive mechanisms, environmental factors, hermaphroditism, and unisexuality. Here we use a de novo genome assembly of New Zealand silver trevally (Pseudocaranx georgianus) together with whole genome sequencing to detect sexually divergent regions, identify candidate genes and develop molecular makers. ResultsThe de novo assembly of an unsexed trevally (Trevally_v1) resulted in an assembly of 579.4 Mb in length, with a N50 of 25.2 Mb. Of the assembled scaffolds, 24 were of chromosome scale, ranging from 11 to 31 Mb. A total of 28416 genes were annotated after 12.8% of the assembly was masked with repetitive elements. Whole genome re-sequencing of 13 sexed trevally (7 males, 6 females) identified sexually divergent regions located on two scaffolds, including a 6 kb region at the proximal end of chromosome 21. Blast analyses revealed similarity between one region and the aromatase genes cyp19 (a1a/b). Males contained higher numbers of heterozygous variants in both regions, while females showed regions of very low read-depth, indicative of deletions. Molecular markers tested on 96 histologically-sexed fish (42 males, 54 females). Three markers amplified in absolute correspondence with sex. ConclusionsThe higher number of heterozygous variants in males combined with deletions in females support a XY sex-determination model, indicating the trevally_v1 genome assembly was based on a male. This sex system contrasts with the ZW-type sex system documented in closely related species. Our results indicate a likely sex-determining function of the cyp19b-like gene, suggesting the molecular pathway of sex determination is somewhat conserved in this family. Our genomic resources will facilitate future comparative genomics works in teleost species, and enable improved insights into the varied sex determination pathways in this group of vertebrates. The sex marker will be a valuable resource for aquaculture breeding programmes, and for determining sex ratios and sex-specific impacts in wild fisheries stocks of this species.

genomics

Genome sequence of Pseudopithomyces chartarum, causal agent of facial eczema (pithomycotoxicosis) in ruminants, and identification of the putative sporidesmin toxin gene cluster

Facial eczema (FE) in grazing ruminants is a debilitating liver syndrome induced by ingestion of sporidesmin, a toxin belonging to the epipolythiodioxopiperazine class of compounds. Sporidesmin is produced in spores of the fungus Pseudopithomyces chartarum, a microbe which colonises leaf litter in pastures. New Zealand has a high occurrence of FE in comparison to other countries as animals are fed predominantly on ryegrass, a species that supports high levels of Pse. chartarum spores. The climate is also particularly conducive for Pse. chartarum growth. Here, we present the genome of Pse. chartarum and identify the putative sporidesmin gene cluster. The Pse. chartarum genome was sequenced using single molecule real-time sequencing (PacBio) and gene models identified. Loci containing genes with homology to the aspirochlorine, sirodesmin PL and gliotoxin cluster genes of Aspergillus oryzae, Leptosphaeria maculans and Aspergillus fumigatus, respectively, were identified by tBLASTn. We identified and annotated an epipolythiodioxopiperazine cluster at a single locus with all the functionality required to synthesise sporidesmin. HighlightsO_LIThe whole genome of Pseudopithomyces chartarum has been sequenced and assembled. C_LIO_LIThe genome is 39.13 Mb, 99% complete, and contains 11,711 protein coding genes. C_LIO_LIA putative sporidesmin A toxin (cause of facial eczema) gene cluster is described. C_LIO_LIThe genomes of Pse. chartarum and the Leptosphaerulina chartarum teleomorph differ. C_LIO_LIComparative genomics is required to further resolve the Pseudopithomyces clade. C_LI

genomics

The Taxus genome provides insights into paclitaxel biosynthesis

The ancient gymnosperm genus Taxus is the exclusive source of the anticancer drug paclitaxel, yet no reference genome sequences are available for comprehensively elucidating the paclitaxel biosynthesis pathway. We have completed a chromosome-level genome of Taxus chinensis var. mairei with a total length of 10.23 Gb. Taxus shared an ancestral whole-genome duplication with the coniferophyte lineage and underwent distinct transposon evolution. We discovered a unique physical and functional grouping of CYP725As (cytochrome P450) in the Taxus genome for paclitaxel biosynthesis. We also identified a gene cluster in the taxadiene biosynthesis, which was mainly formed by gene duplications. This study will facilitate the elucidation of paclitaxel biosynthesis and unleash the biotechnological potential of Taxus. One Sentence SummaryA chromosome-level genome assembly of Taxus chinensis var. mairei uncovers its unique genome evolution process and genetic architectures for the paclitaxel biosynthesis pathway.

genomics

Improved Apis mellifera reference genome based on the alternative long-read-based assemblies

Apis mellifera L., the western honey bee is a major crop pollinator that plays a key role in beekeeping and serves as an important model organism in social behavior studies. Recent efforts have improved on the quality of the honey bee reference genome and developed a chromosome-level assembly of sixteen chromosomes, two of which are gapless. However, the rest suffer from 51 gaps, 160 unplaced/unlocalized scaffolds, and the lack of 2 distal telomeres. The gaps are located at the hard-to-assemble extended highly repetitive chromosomal regions that may contain functional genomic elements. Here, we use de-novo re-assemblies from the most recent reference genome Amel_HAv_3.1 raw reads and other long-read-based assemblies (INRA_AMelMel_1.0, ASM1384120v1, and ASM1384124v1) of the honey bee genome to resolve 13 gaps, five unplaced/unlocalized scaffolds and, the lacking telomeres of the Amel_HAv_3.1. The total length of the resolved gaps is 848,747 bp. The accuracy of the corrected assembly was validated by mapping PacBio reads and performing gene annotation assessment. Comparative analysis suggests that the PacBio-reads-based assemblies of the honey bee genomes failed in the same highly repetitive extended regions of the chromosomes, especially on chromosome 10. To fully resolve these extended repetitive regions, further work using ultra-long Nanopore sequencing would be needed. Our updated assembly facilitates more accurate reference-guided scaffolding and marker/sequence mapping in honey bee genomics studies.

genomics

Chromosome-scale and haplotype-resolved genome assembly of a tetraploid potato cultivar

Potato is the most important tuber crop in the world. However, separate reconstruction of the four haplotypes of its autotetraploid genome remained an unsolved challenge. Here, we report the 3.1 Gb haplotype-resolved (at 99.6% precision), chromosome-scale assembly of the potato cultivar Otava using high-quality long reads coupled with single-cell sequencing of 717 pollen genomes and Hi-C data. Unexpectedly, almost 50% of the genome were found to be identical-by-descent due to recent inbreeding, which contrasted by highly abundant structural rearrangements involving around 20% of the genome. Among 38,214 genes, only 54% were present in four haplotypes with an average of 3.2 copies per gene. Analyzing the leaf transcriptome as example, we found that 11% of the genes featured differently expressed alleles in at least one of the haplotypes, of which 25% are likely regulated through allele-specific DNA methylation. Our work sheds light on the recent breeding history of potato, the functional organization of its tetraploid genome and has the potential to strengthen the future of genomics-assisted breeding.

genomics

Whole human genome 5'-mC methylation analysis using long read nanopore sequencing

DNA methylation is a type of epigenetic modification that affects gene expression regulation and is associated with several human diseases. Microarray and short read sequencing technologies are often used to study 5-methylcytosine (5-mC) modification of CpG dinucleotides in the human genome. Although both technologies produce trustable results, the evaluation of the methylation status of CpG sites suffers from the potential side effects of DNA modification by bisulfite and the ambiguity of mapping short reads in repetitive and highly homologous genomic regions, respectively. Nanopore sequencing is an attractive alternative for the study of 5-mC since the long reads produced by this technology allow to resolve those genomic regions more easily. Moreover, it allows direct sequencing of native DNA molecules using a fast library preparation procedure. In this work we show that 10X coverage depth nanopore sequencing, using DNA from a human cell line, produces 5-mC methylation frequencies consistent with those obtained by methylation microarray and digital restriction enzyme analysis of methylation. In particular, the correlation of methylation values ranged from 0.73 to 0.90 using an average genome sequencing coverage depth <2X or a minimum read support of 17X for each CpG site, respectively. We also showed that a minimum of 5 reads per CpG yields strong correlations (>0.89) between sequencing runs and an almost uniform variation in methylation frequencies of CpGs across the entire value range. Furthermore, nanopore sequencing was able to correctly display methylation frequency patterns according to genomic annotations, including a majority of unmethylated and methylated sites in the CpG islands and inter-CpG island regions, respectively. These results demonstrate that low coverage depth nanopore sequencing is a fast, reliable and unbiased approach to the study of 5-mC in the human genome.

genomics

Segmental duplications and their variation in a complete human genome

Despite their importance in disease and evolution, highly identical segmental duplications (SDs) have been among the last regions of the human reference genome (GRCh38) to be finished. Based on a complete telomere-to-telomere human genome (T2T-CHM13), we present the first comprehensive view of human SD organization. SDs account for nearly one-third of the additional sequence increasing the genome-wide estimate from 5.4% to 7.0% (218 Mbp). An analysis of 266 human genomes shows that 91% of the new T2T-CHM13 SD sequence (68.3 Mbp) better represents human copy number. We find that SDs show increased single-nucleotide variation diversity when compared to unique regions; we characterize methylation signatures that correlate with duplicate gene transcription and predict 182 novel protein-coding gene candidates. We find that 63% (35.11/55.7 Mbp) of acrocentric chromosomes consist of SDs distinct from rDNA and satellite sequences. Acrocentric SDs are 1.75-fold longer (p=0.00034) than other SDs, are frequently shared with autosomal pericentromeric regions, and are heteromorphic among human chromosomes. Comparing long-read assemblies from other human (n=12) and nonhuman primate (n=5) genomes, we use the T2T-CHM13 genome to systematically reconstruct the evolution and structural haplotype diversity of biomedically relevant (LPA, SMN) and duplicated genes (TBC1D3, SRGAP2C, ARHGAP11B) important in the expansion of the human frontal cortex. The analysis reveals unprecedented patterns of structural heterozygosity and massive evolutionary differences in SD organization between humans and their closest living relatives.

genomics

Genomic Abelian Finite Groups

Experimental studies reveal that genome architecture splits into DNA sequence domains suggesting a well-structured genomic architecture, where, for each species, genome populations are integrated by individual mutational variants. Herein, we show that, consistent with the fundamental theorem of Abelian finite groups, the architecture of population genomes from the same or closed related species can be quantitatively represented in terms of the direct sum of homocyclic Abelian groups of prime-power order defined on the genetic code and on the set of DNA bases, where populations can be stratified into subpopulations with the same canonical decomposition into p-groups. Through concrete examples we show that the architectures of current annotated genomic regions including (but not limited to) transcription factors binding-motif, promoter regulatory boxes, exon and intron arrangement associated to gene splicing are subjects for feasible modeling as decomposable Abelian p-groups. Moreover, we show that the epigenomic variations induced by diseases or environmental changes also can be represented as an Abelian group decomposable into homocyclic Abelian p-groups. The nexus between the direct sum of homocycle Abelian p-groups and the endomorphism ring paved the ways to unveil unsuspected stochastic-deterministic logical propositions ruling the ensemble of genomic regions. Our study aims to set the basis for concrete applications of the theory in computational biology and bioinformatics. Consistently with this goal, a computational tool designed for the analysis of fixed mutational events in gene/genome populations represented as endomorphisms and automorphisms is provided. Results suggest that complex local architectures and evolutionary features no evident through the direct experimentation can be unveiled through the analysis of the endomorphism ring and the subsequent application of machine learning approaches for the identification of stochastic-deterministic logical rules (reflecting the evolutionary pressure on the region) constraining the set of possible mutational events (represented as homomorphisms) and the evolutionary paths.

genomics

High-quality Arabidopsis thaliana genome assembly with Nanopore and HiFi long reads

Arabidopsis thaliana is an important and long-established model species for plant molecular biology, genetics, epigenetics, and genomics. However, the latest version of reference genome still contains significant number of missing segments. Here, we report a high-quality and almost complete Col-0 genome assembly with two gaps (Col-XJTU) using combination of Oxford Nanopore Technology ultra-long reads, PacBio high-fidelity long reads, and Hi-C data. The total genome assembly size is 133,725,193 bp, introducing 14.6 Mb of novel sequences compared to the TAIR10.1 reference genome. All five chromosomes of Col-XJTU assembly are highly accurate with consensus quality (QV) scores > 60 (ranging from 62 to 68), which are higher than those of TAIR10.1 reference (QV scores ranging from 45 to 52). We have completely resolved chromosome (Chr) 3 and Chr5 in a telomere-to-telomere manner. Chr4 has been completely resolved except the nucleolar organizing regions, which comprise long repetitive DNA fragments. The Chr1 centromere (CEN1), reportedly around 9 Mb in length, is particularly challenging to assemble due to the presence of tens of thousands of CEN180 satellite repeats. Using the cutting-edge sequencing data and novel computational approaches, we assembled about 4 Mb of sequence for CEN1 and a 3.5-Mb-long CEN2. We investigated the structure and epigenetics of centromeres. We detected four clusters of CEN180 monomers, and found that the centromere-specific histone H3-like protein (CENH3) exhibits a strong preference for CEN180 cluster 3. Moreover, we observed hypomethylation patterns in CENH3-enriched regions. We believe that this high-quality genome assembly, Col-XJTU, would serve as a valuable reference to better understand the global pattern of centromeric polymorphisms, as well as genetic and epigenetic features in plants.

genomics

Quantitative Analysis of Genomic Sequences of Virus RNAs Using a Metric-Based Algorithm

This work aims to study the virus RNAs using a novel algorithm for accelerated exploring any-length genomic fragments in sequences using Hamming distance between the binary-expressed characters of an RNA and query patterns. The found repetitive genomic sub-sequences of different lengths were placed on one plot as genomic trajectories (walks) to increase the effectiveness of geometrical multi-scale genomic studies. Primary attention was paid to the building and analysis of the atg-triplet walks composing the schemes or skeletons of the viral RNAs. The 1-D distributions of these codon-starting atg-triplets were built with the single-symbol walks for full-scale analyses. The visual examination was followed by calculating statistical parameters of genomic sequences, including the estimation of geometry deviation and fractal properties of inter-atg distances. This approach was applied to the SARS CoV-2, MERS CoV, Dengue and Ebola viruses, whose complete genomic sequences are taken from GenBank and GISAID databases. The relative stability of these distributions for SARS CoV-2 and MERS CoV viruses was found, unlike the Dengue and Ebola distributions that showed an increased deviation of their geometrical and fractal characteristics of atg-distributions. The results of this work can found in classification of the virus families and in the study of their mutation.

genomics