bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

A chromosome level genome of Astyanax mexicanus surface fish for comparing population-specific genetic differences contributing to trait evolution.

Identifying the genetic factors that underlie complex traits is central to understanding the mechanistic underpinnings of evolution. In nature, adaptation to severe environmental change, such as encountered following colonization of caves, has dramatically altered genomes of species over varied time spans. Genomic sequencing approaches have identified mutations associated with troglomorphic trait evolution, but the functional impacts of these mutations remain poorly understood. The Mexican Tetra, Astyanax mexicanus, is abundant in the surface waters of northeastern Mexico, and also inhabits at least 30 different caves in the region. Cave-dwelling A. mexicanus morphs are well adapted to subterranean life and many populations appear to have evolved troglomorphic traits independently, while the surface-dwelling populations can be used as a proxy for the ancestral form. Here we present a high-resolution, chromosome-level surface fish genome, enabling the first genome-wide comparison between surface fish and cavefish populations. Using this resource, we performed quantitative trait locus (QTL) mapping analyses for pigmentation and eye size and found new candidate genes for eye loss such as dusp26. We used CRISPR gene editing in A. mexicanus to confirm the essential role of a gene within an eye size QTL, rx3, in eye formation. We also generated the first genome-wide evaluation of deletion variability that includes an analysis of the impact on protein-coding genes across cavefish populations to gain insight into this potential source of cave adaptation. The new surface fish genome reference now provides a more complete resource for comparative, functional, developmental and genetic studies of drastic trait differences within a species.Competing Interest StatementThe authors have declared no competing interest.View Full Text

genomics

Near-infrared spectroscopy outperforms genomic selection for predicting sugarcane feedstock quality traits

The main objectives of this study were to evaluate the prediction performance of genomic and near-infrared spectroscopy (NIR) data and whether the integration of genomic and NIR predictor variables can increase the prediction accuracy of two feedstock quality traits (fiber and sucrose content) in a sugarcane population (Saccharum spp.). The following three modeling strategies were compared: M1 (genome-based prediction), M2 (NIR-based prediction), and M3 (integration of genomics s and NIR wavenumbers). Data were collected from a commercial population comprised of three hundred and eighty-five individuals, genotyped for single nucleotide polymorphisms (SNPs) and screened using NIR spectroscopy. We compared partial least squares (PLS) and BayesB regression methods to estimate marker and wavenumber effects. In order to assess model performance, we employed random sub-sampling cross-validation to calculate the mean Pearson correlation coefficient between observed and predicted genotypic values. Our results showed that models fitted using BayesB were most predictive than PLS models. We found that NIR (M2) provided the highest prediction accuracy, whereas genomics (M1) presented the lowest predictive ability, regardless of the measured traits and regression methods used. The integration of predictors derived from NIR spectroscopy and genomics into a single model (M3) did not significantly improve the prediction accuracy for the two traits evaluated. These findings suggest that NIR-based prediction can be an effective strategy for predicting the genotypic value of sugarcane clones.

genomics

The impact of bottlenecks and inbreeding on the genome of the endangered Pyrenean desman

The Pyrenean desman (Galemys pyrenaicus) is a small semiaquatic mammal endemic to the Iberian Peninsula. Despite its limited range, this species presents a strong genetic structure due to past isolation in glacial refugia and subsequent bottlenecks. Additionally, some populations are highly fragmented today as a consequence of river barriers, causing substantial levels of inbreeding. These features make the Pyrenean desman a unique model in which to study the genomic footprints of differentiation, bottlenecks and extreme isolation in an endangered species. The complete genome of the Pyrenean desman was assembled using a Bloom filter-based approach. An analysis of the 1.83 Gb reference genome and the sequencing of five additional individuals from different evolutionary units allowed us to detect its main genomic characteristics. The population differentiation of the species was reflected in highly distinctive demographic trajectories. A severe population bottleneck during the postglacial recolonization of the eastern Pyrenees created the lowest genomic heterozygosity ever recorded in a mammal. Moreover, isolation and inbreeding gave rise to a high proportion of runs of homozygosity (ROH). Despite these extremely low levels of genetic diversity, two key multigene families from an eco-evolutionary perspective that need to be genetically variable, the major histocompatibility complex and olfactory receptor genes, showed heterozygosity excess in the majority of individuals. Furthermore, these two classes of genes were significantly less abundant than expected within ROH. These results allow us to characterize important genomic health indicators for each individual, information that may be crucial for the conservation and management of the species.

genomics

Phylogenomic analysis of SARS-CoV-2 genomes from western India reveals unique linked mutations

India has become the third worst-hit nation by the COVID-19 pandemic caused by the SARS-CoV-2 virus. Here, we investigated the molecular, phylogenomic, and evolutionary dynamics of SARS-CoV-2 in western India, the most affected region of the country. A total of 90 genomes were sequenced. Four nucleotide variants, namely C241T, C3037T, C14408T (Pro4715Leu), and A23403G (Asp614Gly), located at 5UTR, Orf1a, Orf1b, and Spike protein regions of the genome, respectively, were predominant and ubiquitous (90%). Phylogenetic analysis of the genomes revealed four distinct clusters, formed owing to different variants. The major cluster (cluster 4) is distinguished by mutations C313T, C5700A, G28881A are unique patterns and observed in 45% of samples. We thus report a newly emerging pattern of linked mutations. The predominance of these linked mutations suggests that they are likely a part of the viral fitness landscape. A novel and distinct pattern of mutations in the viral strains of each of the districts was observed. The Satara district viral strains showed mutations primarily at the 3' end of the genome, while Nashik district viral strains displayed mutations at the 5' end of the genome. Characterization of Pune strains showed that a novel variant has overtaken the other strains. Examination of the frequency of three mutations i.e., C313T, C5700A, G28881A in symptomatic versus asymptomatic patients indicated an increased occurrence in symptomatic cases, which is more prominent in females. The age-wise specific pattern of mutation is observed. Mutations C18877T, G20326A, G24794T, G25563T, G26152T, and C26735T are found in more than 30% study samples in the age group of 10-25. Intriguingly, these mutations are not detected in the higher age range 61-80. These findings portray the prevalence of unique linked mutations in SARS-CoV-2 in western India and their prevalence in symptomatic patients. ImportanceElucidation of the SARS-CoV-2 mutational landscape within a specific geographical location, and its relationship with age and symptoms, is essential to understand its local transmission dynamics and control. Here we present the first comprehensive study on genome and mutation pattern analysis of SARS-CoV-2 from the western part of India, the worst affected region by the pandemic. Our analysis revealed three unique linked mutations, which are prevalent in most of the sequences studied. These may serve as a molecular marker to track the spread of this viral variant to different places.

genomics

The genome of low-chill Chinese plum 'Sanyueli' (Prunus salicina Lindl.) provides insights into the regulation of chilling requirement of flower bud

Chinese plum (Prunus salicina Lindl.) is a stone fruit that belongs to the Prunus genus and plays an important role in the global production of plum. In this study, we report the genome sequence of the Chinese plum Sanyueli, which is known to have a low-chill requirement for flower bud break. The assembled genome size was 308.06 Mb, with a contig N50 of 815.7 kb. A total of 30,159 protein-coding genes were predicted from the genome and 56.4% (173.39 Mb) of the genome was annotated as repetitive sequence. Bud dormancy is influenced by chilling requirement in plum and partly controlled by DORMANCY ASSOCIATED MADS-box (DAM) genes. Six tandemly arrayed PsDAM genes were identified in the assembled genome. Sequence analysis of PsDAM6 in Sanyuelirevealed the presence of large insertions in the intron and exon regions. Transcriptome analysis indicated that the expression of PsDAM6 in the dormant flower buds of Sanyueli was significantly lower than that in the dormant flower buds of the high chill requiring Furongli plum. In addition, the expression of PsDAM6 was repressed by chilling treatment. The genome sequence of Sanyueli plum provides a valuable resource for elucidating the molecular mechanisms responsible for the regulation of chilling requirements, and is also useful for the identification of the genes involved in the control of other important agronomic traits and molecular breeding in plum.

genomics

Genome wide efficiency profiling reveals modulation of maintenance and de novo methylation by Tets

A precise understanding of DNA methylation dynamics on a genome wide scale is of great importance for the comprehensive investigation of a variety of biological processes such as reprogramming of somatic cells to iPSCs, cell differentiation and also cancer development. To date, a complex integration of multiple and distinct genome wide data sets is required to derive the global activity of DNA modifying enzymes. We present GwEEP - Genome-wide Epigenetic Efficiency Profiling as a versatile approach to infer dynamic efficiency changes of DNA modifying enzymes at base pair resolution on a genome wide scale. GwEEP relies on genome wide oxidative Hairpin Bisulfite sequencing (HPoxBS) data sets, which are translated by a sophisticated hidden Markov model into quantitative enzyme efficiencies with reported confidence around the estimates. GwEEP in its present form predicts de novo and maintenance methylation efficiencies of Dnmts, as well as the hydroxylation efficiency of Tets but its purposefully flexible design allows to capture further oxidation processes such as formylation and carboxylation given available data in the future. Applied to a well characterized ES cell model, GwEEP precisely predicts the complex epigenetic changes following a Serum-to-2i shift i.e., (i) instant reduction in maintenance efficiency (ii) gradually decreasing de novo methylation efficiency and (iii) increasing Tet efficiencies. In addition, a complementary analysis of Tet triple knock-out ES cells confirms the previous hypothesized mutual interference of Dnmts and Tets. GwEEP is applicable to a wide range of biological samples including cell lines, but also tissues and primary cell types. MOTIVATIONDynamic changes of DNA methylation patterns are a common phenomenon in epigenetics. Although a stable DNA methylation profile is essential for cell identity, developmental processes require the rearrangement of 5-methylcytosine in the genome. Stable methylation patterns are the result of balanced Dnmts and Tets activities, while methylome transformation results from a coordinated change in Dnmt and Tet efficiencies. Such transformations occur on a global scale, for example during the reprogramming of maternal and paternal methylation patterns and the establishment of novel cell type specific methylomes during embryonic development in vivo, but also in vitro during (re)programming of induced pluripotent stem cells, as well as somatic cells. In addition, local (de)methylation events are key for gene regulation during cell differentiation. A detailed understanding of Dnmt and Tet cooperation is essential for understanding natural epigenetic adaptation as well as optimization of in vitro (re)programming protocols. For this purpose, we developed a pipeline for quantitative and precise estimation of Dnmt and Tet activity. Using only double strand methylation information, GwEEP infers accurate maintenance and de novo methylation efficiency of Dnmts, as well as hydroxylation efficiency of Tets at single base resolution. Thus, we believe GwEEP provides a powerful tool for the investigation of methylome rearrangements in various systems.

genomics

Rapid cost-effective viral genome sequencing by V-seq

Conventional methods for viral genome sequencing largely use metatranscriptomic approaches or, alternatively, enrich for viral genomes by amplicon sequencing with virus-specific PCR or hybridization-based capture. These existing methods are costly, require extensive sample handling time, and have limited throughput. Here, we describe V-seq, an inexpensive, fast, and scalable method that performs targeted viral genome sequencing by multiplexing virus-specific primers at the cDNA synthesis step. We designed densely tiled reverse transcription (RT) primers across the SARS-CoV-2 genome, with a subset of hexamers at the 3 end to minimize mis-priming from the abundant human rRNA repeats and human RNA PolII transcriptome. We found that overlapping RT primers do not interfere, but rather act in concert to improve viral genome coverage in samples with low viral load. We provide a path to optimize V-seq with SARS-CoV-2 as an example. We anticipate that V-seq can be applied to investigate genome evolution and track outbreaks of RNA viruses in a cost-effective manner. More broadly, the multiplexed RT approach by V-seq can be generalized to other applications of targeted RNA sequencing.

genomics

Genomic selection for any dairy breeding program via optimized investment in phenotyping and genotyping

This paper evaluates the potential of maximizing genetic gain in dairy cattle breeding by optimizing investment into phenotyping and genotyping. Conventional breeding focuses on phenotyping selection candidates or their close relatives to maximize selection accuracy for breeders and quality assurance for producers. Genomic selection decoupled phenotyping and selection and through this increased genetic gain per year compared to the conventional selection. Although genomic selection is established in well-resourced breeding programs, small populations and developing countries still struggle with the implementation. The main issues include the lack of training animals and lack of financial resources. To address this, we simulated a case-study of a small dairy population with a number of scenarios with equal resources yet varied use of resources for phenotyping and genotyping. The conventional progeny testing scenario had 11 phenotype records per lactation. In genomic scenarios, we reduced phenotyping to between 10 and 1 phenotype records per lactation and invested the saved resources into genotyping. We tested these scenarios at different relative prices of phenotyping to genotyping and with or without an initial training population for genomic selection. Reallocating a part of phenotyping resources for repeated milk records to genotyping increased genetic gain compared to the conventional scenario regardless of the amount and relative cost of phenotyping, and the availability of an initial training population. Genetic gain increased by increasing genotyping, despite reduced phenotyping. High-genotyping scenarios even saved resources. Genomic scenarios expectedly increased accuracy for young non-phenotyped male and female candidates, but also cows. This study shows that breeding programs should optimize investment into phenotyping and genotyping to maximise return on investment. Our results suggest that any dairy breeding program using conventional progeny testing with repeated milk records can implement genomic selection without increasing the level of investment.

genomics

Improvements to the ARTIC multiplex PCR method for SARS-CoV-2 genome sequencing using nanopore

Genome sequencing has been widely deployed to study the evolution of SARS-CoV-2 with more than 90,000 genome sequences uploaded to the GISAID database. We published a method for SARS-CoV-2 genome sequencing (https://www.protocols.io/view/ncov-2019-sequencing-protocol-bbmuik6w) online on January 22, 2020. This approach has rapidly become the most popular method for sequencing SARS-CoV-2 due to its simplicity and cost-effectiveness. Here we present improvements to the original protocol: i) an updated primer scheme with 22 additional primers to improve genome coverage, ii) a streamlined library preparation workflow which improves demultiplexing rate for up to 96 samples and reduces hands-on time by several hours and iii) cost savings which bring the reagent cost down to {pound}10 per sample making it practical for individual labs to sequence thousands of SARS-CoV-2 genomes to support national and international genomic epidemiology efforts.

genomics

Genome sequencing of turmeric provides evolutionary insights into its medicinal properties

Curcuma longa, or turmeric, is traditionally known for its immense medicinal properties and has diverse therapeutic applications. However, the absence of a reference genome sequence is a limiting factor in understanding the genomic basis of the origin of its medicinal properties. In this study, we present the draft genome sequence of Curcuma longa, the first species sequenced from Zingiberaceae plant family, constructed using 10x Genomics linked reads. For comprehensive gene set prediction and for insights into its gene expression, the transcriptome sequencing of leaf tissue was also performed. The draft genome assembly had a size of 1.24 Gbp with ~74% repetitive sequences, and contained 56,036 coding gene sequences. The phylogenetic position of Curcuma longa was resolved through a comprehensive genome-wide phylogenetic analysis with 16 other plant species. Using 5,294 orthogroups, the comparative evolutionary analysis performed across 17 species including Curcuma longa revealed evolution in genes associated with secondary metabolism, plant phytohormones signaling, and various biotic and abiotic stress tolerance responses. These mechanisms are crucial for perennial and rhizomatous plants such as Curcuma longa for defense and environmental stress tolerance via production of secondary metabolites, which are associated with the wide range of medicinal properties in Curcuma longa.

genomics

Chromosome-level genome assembly of a benthic associated Syngnathiformes species: the common dragonet, Callionymus lyra

BackgroundThe common dragonet, Callionymus lyra, is one of three Callionymus species inhabiting the North Sea. All three species show strong sexual dimorphism. The males show strong morphological differentiation, e.g., species-specific colouration and size relations, while the females of different species have few distinguishing characters. Callionymus belongs to the benthic associated clade of the order Syngnathiformes. The benthic associated clade so far is not represented by genome data and serves as an important outgroup to understand the morphological transformation in long-snouted syngnatiforms such as seahorses and pipefishes. FindingsHere, we present the chromosome-level genome assembly of C. lyra. We applied Oxford Nanopore Technologies long-read sequencing, short-read DNBseq, and proximity-ligation-based scaffolding to generate a high-quality genome assembly. The resulting assembly has a contig N50 of 2.2 Mbp, a scaffold N50 of 26.7 Mbp. The total assembly length is 568.7 Mbp, of which over 538 Mbp were scaffolded into 19 chromosome-length scaffolds. The identification of 94.5% of complete BUSCO genes indicates high assembly completeness. Additionally, we sequenced and assembled a multi-tissue transcriptome with a total length of 255.5 Mbp that was used to aid the annotation of the genome assembly. The annotation resulted in 19,849 annotated transcripts and identified a repeat content of 27.66%. ConclusionsThe chromosome-level assembly of C. lyra provides a high-quality reference genome for future population genomic, phylogenomic, and phylogeographic analyses.

genomics

Whole-genome sequencing of 1,171 elderly admixed individuals from the largest Latin American metropolis (Sao Paulo, Brazil)

As whole-genome sequencing (WGS) becomes the gold standard tool for studying population genomics and medical applications, data on diverse non-European and admixed individuals are still scarce. Here, we present a high-coverage WGS dataset of 1,171 highly admixed elderly Brazilians from a census-based cohort, providing over 76 million variants, of which ~2 million are absent from large public databases. WGS enabled identifying ~2,000 novel mobile element insertions, nearly 5Mb of genomic segments absent from human genome reference, and over 140 novel alleles from HLA genes. We reclassified and curated nearly four hundred variant's pathogenicity assertions in genes associated with dominantly inherited Mendelian disorders and calculated the incidence for selected recessive disorders, demonstrating the clinical usefulness of the present study. Finally, we observed that whole-genome and HLA imputation could be significantly improved compared to available datasets since rare variation represents the largest proportion of input from WGS. These results demonstrate that even smaller sample sizes of underrepresented populations bring relevant data for genomic studies, especially when exploring analyses allowed only by WGS.

genomics

A comparative genomics and immunoinformatics approach to identify epitope-based peptide vaccine candidates against bovine hemoplasmosis

Mycoplasma wenyonii and Candidatus Mycoplasma haemobos have been described as major hemoplasmas that infect cattle worldwide. Currently, three bovine hemoplasma genomes are known. The aim of this work was to know the main genomic characteristics and the evolutionary relationships between hemoplasmas, as well as to provide a list of epitopes identified by immunoinformatics that could be used as vaccine candidates against bovine hemoplasmosis. So far, there is not a vaccine to prevent this disease that impact economically in cattle production around the world. In this work, we used comparative genomics to analyze the genomes of the hemoplasmas so far reported. As a result, we confirm that Ca. M haemobos INIFAP01 is a divergent species from M. wenyonii INIFAP02 and M. wenyonii Massachusetts. Although both strains of M. wenyonii have genomes with similar characteristics (length, G+C content, tRNAs and position of rRNAs) they have different structures (alignment coverage and identity of 51.58 and 79.37%, respectively). The correct genomic characterization of bovine hemoplasmas, never studied before, will allow to develop better molecular detection methods, to understand the possible pathogenic mechanisms of these bacteria and to identify epitopes sequences that could be used in the vaccine design.

genomics

Genomic sequencing confirms absence of introgression despite past hybridisation in a critically endangered bird

Genetic swamping resulting from interspecific hybridisation can increase extinction risk for threatened species. The development of high-throughput and reduced-representation genomic sequencing and analyses to generate large numbers of high resolution genomic markers has the potential to reveal introgression previously undetected using small numbers of genetic markers. However, few studies to date have implemented genomic tools to assess the extent of interspecific hybridisation in threatened species. Here we investigate the utility of genome-wide single nucleotide polymorphisms (SNPs) to detect introgression resulting from past interspecific hybridisation in one of the worlds rarest birds. Anthropogenic impacts have resulted in hybridisation and subsequent backcrossing of the critically endangered Aotearoa New Zealand endemic kak[i] (black stilts; Himantopus novaezelandiae) with the non-threatened self-introduced congeneric poaka (Aotearoa New Zealand population of pied stilts, Himantopus himantopus leucocephalus), yet genetic analyses with a limited set of microsatellite markers revealed no evidence of introgression of poaka genetic material in kak[i], excluding one individual. We use genomic data for [~]63% of the wild adult kak[i] population to reassess the extent of introgression resulting from hybridisation between kak[i] and poaka. Consistent with previous genetic analyses, we detected no introgression from poaka into kak[i]. These collective results indicate that, for kak[i], existing microsatellite markers provide a robust, cost-effective approach to detect cryptic hybrids. Further, for well-differentiated species, the use of genomic markers may not be required to detect admixed individuals.

genomics

De novo structural mutation rates and gamete-of-origin biases revealed through genome sequencing of 2,396 families

Each human genome includes de novo mutations that arose during gametogenesis. While these germline mutations represent a fundamental source of new genetic diversity, they can also create deleterious alleles that impact fitness. The germline mutation rate for single nucleotide variants and factors that significantly influence this rate, such as parental age, are now well established. However, far less is known about the frequency, distribution, and features that impact de novo structural mutations. We report a large, family-based study of germline mutations, excluding aneuploidy, that affect genome structure among 572 genomes from 33 families in a multigenerational CEPH-Utah cohort and 2,363 cases of non-familial autism spectrum disorder (ASD), 1,938 unaffected siblings, and both parents (9,599 genomes in total). We find that de novo structural mutations detected by alignment-based, short-read WGS occurred at an overall rate of at least 0.160 events per genome in unaffected individuals and was significantly higher (0.206 per genome) in ASD cases. In both probands and unaffected samples, nearly 73% of de novo structural mutations arose in paternal gametes, and predict most de novo structural mutations to be caused by mutational mechanisms that do not require sequence homology. After multiple testing correction we did not observe a statistically significant correlation between parental age and the rate of de novo structural variation in offspring. These results highlight that a spectrum of mutational mechanisms contribute to germline structural mutations, and that these mechanisms likely have markedly different rates and selective pressures than those leading to point mutations.

genomics

The Thermosynechococcus genus: wide environmental distribution, but a highly conserved genomic core

Cyanobacteria thrive in very diverse environments. However, questions remain about possible growth limitations in ancient environmental conditions. As a single genus, the Thermosynechococcus are cosmopolitan and live in chemically diverse habitats. To understand the genetic basis for this, we compared the protein coding component of Thermosynechococcus genomes. Supplementing the known genetic diversity of Thermosynechococcus, we report draft metagenome-assembled genomes of two Thermosynechococcus recovered from ferrous carbonate hot springs in Japan. We find that as a genus, Thermosynechococcus is genomically conserved, having a small pan-genome with few accessory genes per individual strain and only 14 putative orthologous protein groups appearing in all Thermosynechococcus but not in any other cyanobacteria in our analysis. Furthermore, by comparing orthologous protein groups, including an analysis of genes encoding proteins with an iron related function (uptake, storage or utilization), no clear differences in genetic content, or adaptive mechanisms could be detected between genus members, despite the range of environments they inhabit. Overall, our results highlight a seemingly innate ability for Thermosynechococcus to inhabit diverse habitats without having undergone substantial genomic adaptation to accommodate this. The finding of Thermosynechococcus in both hot and high iron environments without adaptation recognizable from the perspective of the proteome has implications for understanding the basis of thermophily within this clade, and also for understanding the possible genetic basis for high iron tolerance in cyanobacteria on early Earth. The conserved core genome may be indicative of an allopatric lifestyle - or reduced genetic complexity of hot spring habitats relative to other environments.

genomics

Machine-learning predicts genomic determinants of meiosis-driven structural variation in a eukaryotic pathogen

Species harbor extensive structural variation underpinning recent adaptive evolution and major disease phenotypes. Most sequence rearrangements are generated non-randomly along the genome through non-allelic recombination and transposable element activity. However, the causality between genomic features and the induction of new rearrangements is poorly established. Here, we analyze a global set of telomere-to-telomere genome assemblies of a major fungal pathogen of wheat to establish a nucleotide-level map of structural variation. We show that the recent emergence of pesticide resistance has been disproportionally driven by rearrangements. We used machine-learning to train a model on structural variation events based on 30 chromosomal sequence features. We show that base composition and gene density are the major determinants of structural variation. Low-copy LINE and Gypsy retrotransposons explain most inversion, indel and duplication events. We retrain our model on Arabidopsis thaliana and show that our modelling approach can be extended to more complex genomes. Finally, we analyzed complete genomes of haploid offspring in a four-generation pedigree. Meiotic crossover locations were enriched for newly generated structural variation consistent with crossovers being mutational hotspots. The model trained on species-wide structural variation predicted the position of >74% of the newly generated variants along the pedigree. The predictive power highlights causality between specific sequence features and the induction of chromosomal rearrangements. Our work demonstrates that training sequence-derived models can accurately identify regions of intrinsic DNA instability in eukaryotic genomes.

genomics

Predicting the animal hosts of coronaviruses from compositional biases of spike protein and whole genome sequences through machine learning

The COVID-19 pandemic has demonstrated the serious potential for novel zoonotic coronaviruses to emerge and cause major outbreaks. The immediate animal origin of the causative virus, SARS-CoV-2, remains unknown, a notoriously challenging task for emerging disease investigations. Coevolution with hosts leads to specific evolutionary signatures within viral genomes that can inform likely animal origins. We obtained a set of 650 spike protein and 511 whole genome nucleotide sequences from 225 and 187 viruses belonging to the family Coronaviridae, respectively. We then trained random forest models independently on genome composition biases of spike protein and whole genome sequences, including dinucleotide and codon usage biases in order to predict animal host (of nine possible categories, including human). In hold-one-out cross-validation, predictive accuracy on unseen coronaviruses consistently reached [~]73%, indicating evolutionary signal in spike proteins to be just as informative as whole genome sequences. However, different composition biases were informative in each case. Applying optimised random forest models to classify human sequences of MERS-CoV and SARS-CoV revealed evolutionary signatures consistent with their recognised intermediate hosts (camelids, carnivores), while human sequences of SARS-CoV-2 were predicted as having bat hosts (suborder Yinpterochiroptera), supporting bats as the suspected origins of the current pandemic. In addition to phylogeny, variation in genome composition can act as an informative approach to predict emerging virus traits as soon as sequences are available. More widely, this work demonstrates the potential in combining genetic resources with machine learning algorithms to address long-standing challenges in emerging infectious diseases.

genomics