bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

panRGP: a pangenome-based method to predict genomic islands and explore their diversity

MotivationHorizontal gene transfer (HGT) is a major source of variability in prokaryotic genomes. Regions of Genome Plasticity (RGPs) are clusters of genes located in highly variable genomic regions. Most of them arise from HGT and correspond to Genomic Islands (GIs). The study of those regions at the species level has become increasingly difficult with the data deluge of genomes. To date no methods are available to identify GIs using hundreds of genomes to explore their diversity. ResultsWe present here the panRGP method that predicts RGPs using pangenome graphs made of all available genomes for a given species. It allows the study of thousands of genomes in order to access the diversity of RGPs and to predict spots of insertions. It gave the best predictions when benchmarked along other GI detection tools against a reference dataset. In addition, we illustrated its use on Metagenome Assembled Genomes (MAGs) by redefining the borders of the leuX tRNA hotspot, a well studied spot of insertion in Escherichia coli. panRPG is a scalable and reliable tool to predict GIs and spots making it an ideal approach for large comparative studies. AvailabilityThe methods presented in the current work are available through the following software: https://github.com/labgem/PPanGGOLiN. Detailed results and scripts to compute the benchmark metrics are available at https://github.com/axbazin/panrgp_supdata. Contactvallenet@genoscope.cns.fr and acalteau@genoscope.cns.fr Supplementary informationNone.

bioinformatics↗

Population genomics insights into the recent evolution of SARS-CoV-2

The current coronavirus disease 2019 (COVID-19) pandemic is caused by the SARS-CoV-2 virus and is still spreading rapidly worldwide. Full-genome-sequence computational analysis of the SARS-CoV-2 genome will allow us to understand the recent evolutionary events and adaptability mechanisms more accurately, as there is still neither effective therapeutic nor prophylactic strategy. In this study, we used population genetics analysis to infer the mutation rate and plausible recombination events that may have contributed to the evolution of the SARS-CoV-2 virus. Furthermore, we localized targets of recent and strong positive selection. The genomic regions that appear to be under positive selection are largely co-localized with regions in which recombination from non-human hosts appeared to have taken place in the past. Our results suggest that the pangolin coronavirus genome may have contributed to the SARS-CoV-2 genome by recombination with the bat coronavirus genome. However, we find evidence for additional recombination events that involve coronavirus genomes from other hosts, i.e., Hedgehog and Sparrow. Even though recombination events within human hosts cannot be directly assessed, due to the high similarity of SARS-CoV-2 genomes, we infer that recombinations may have recently occurred within human hosts using a linkage disequilibrium analysis. In addition, we employed an Approximate Bayesian Computation approach to estimate the parameters of a demographic scenario involving an exponential growth of the size of the SARS-CoV-2 populations that have infected European, Asian and Northern American cohorts, and we demonstrated that a rapid exponential growth in population size can support the observed polymorphism patterns in SARS-CoV-2 genomes.

evolutionary biology↗

COVIDier: A Deep-learning Tool For Coronaviruses Genome And Virulence Proteins Classification

COVID-19, caused by SARS-CoV-2 infection, has already reached pandemic proportions in a matter of a few weeks. At the time of writing this manuscript, the unprecedented public health crisis caused more than 2.5 million cases with a mortality range of 5-7%. The SARS-CoV-2, also called novel Coronavirus, is related to both SARS-CoV and bat SARS. Great efforts have been spent to control the pandemic that has become a significant burden on the health systems in a short time. Since the emergence of the crisis, a great number of researchers started to use the AI tools to identify drugs, diagnosing using CT scan images, scanning body temperature, and classifying the severity of the disease. The emergence of variants of the SARS-CoV-2 genome is a challenging problem with expected serious consequences on the management of the disease. Here, we introduce COVIDier, a deep learning-based software that is enabled to classify the different genomes of Alpha coronavirus, Beta coronavirus, MERS, SARS-CoV-1, SARS-CoV-2, and bronchitis-CoV. COVIDier was trained on 1925 genomes, belonging to the three families of SARS retrieved from NCBI Database to propose a new method to train deep learning model trained on genome data using Multi-layer Perceptron Classifier (MLPClassifier), a deep learning algorithm, that could blindly predict the virus family name from the genome of by predicting the statistically similar genome from training data to the given genome. COVIDier able to predict how close the emerging novel genomes of SARS to the known genomes with accuracy 99%. COVIDier can replace tools like BLAST that consume higher CPU and time.

bioinformatics↗

Verification of genetic engineering in yeasts with nanopore whole genome sequencing

Yeast genomes can be assembled from sequencing data, but genome integrations and episomal plasmids often fail to be resolved with accuracy, completeness, and contiguity. Resolution of these features is critical for many synthetic biology applications, including strain quality control and identifying engineering in unknown samples. Here, we report an integrated workflow, named Prymetime, that uses sequencing reads from inexpensive NGS platforms, assembly and error correction software, and a list of synthetic biology parts to achieve accurate whole genome sequences of yeasts with engineering annotated. To build the workflow, we first determined which sequencing methods and software packages returned an accurate, complete, and contiguous genome of an engineered S. cerevisiae strain with two similar plasmids and an integrated pathway. We then developed a sequence feature annotation step that labels synthetic biology parts from a standard list of yeast engineering sequences or from a custom sequence list. We validated the workflow by sequencing a collection of 15 engineered yeasts built from different parent S. cerevisiae and nonconventional yeast strains. We show that each integrated pathway and episomal plasmid can be correctly assembled and annotated, even in strains that have part repeats and multiple similar plasmids. Interestingly, Prymetime was able to identify deletions and unintended integrations that were subsequently confirmed by other methods. Furthermore, the whole genomes are accurate, complete, and contiguous. To illustrate this clearly, we used a publicly available S. cerevisiae CEN.PK113 reference genome and the accompanying reads to show that a Prymetime genome assembly is equivalent to the reference using several standard metrics. Finally, we used Prymetime to resequence the nonconventional yeasts Y. lipolytica Po1f and K. phaffii CBS 7435, producing an improved genome assembly for each strain. Thus, our workflow can achieve accurate, complete, and contiguous whole genome sequences of yeast strains before and after engineering. Therefore, Prymetime enables NGS-based strain quality control through assembly and identification of engineering features.

synthetic biology↗

COVID-Align: Accurate online alignment of hCoV-19 genomes using a profile HMM

MotivationThe first cases of the COVID-19 pandemic emerged in December 2019. Until the end of February 2020, the number of available genomes was below 1,000, and their multiple alignment was easily achieved using standard approaches. Subsequently, the availability of genomes has grown dramatically. Moreover, some genomes are of low quality with sequencing/assembly errors, making accurate re-alignment of all genomes nearly impossible on a daily basis. A more efficient, yet accurate approach was clearly required to pursue all subsequent bioinformatics analyses of this crucial data. ResultshCoV-19 genomes are highly conserved, with very few indels and no recombination. This makes the profile HMM approach particularly well suited to align new genomes, add them to an existing alignment and filter problematic ones. Using a core of [~]2,500 high quality genomes, we estimated a profile using HMMER, and implemented this profile in COVID-Align, a user-friendly interface to be used online or as standalone via Docker. The alignment of 1,000 genomes requires less than 20mn on our cluster. Moreover, COVID-Align provides summary statistics, which can be used to determine the sequencing quality and evolutionary novelty of input genomes (e.g. number of new mutations and indels). Availabilityhttps://covalign.pasteur.cloud, hub.docker.com/r/evolbioinfo/covid-align Contactsolivier.gascuel@pasteur.fr, frederic.lemoine@pasteur.fr Supplementary informationSupplementary information is available at Bioinformatics online.

bioinformatics↗

Characterization of systemic genomic instability in budding yeast

Conventional models of genome evolution are centered around the principle that mutations form independently of each other and build up slowly over time. We characterized the occurrence of bursts of genome-wide loss-of-heterozygosity (LOH) in Saccharomyces cerevisiae, providing support for an additional non-independent and faster mode of mutation accumulation. We initially characterized a yeast clone isolated for carrying an LOH event at a specific chromosome site, and surprisingly, found that it also carried multiple unselected rearrangements elsewhere in its genome. Whole genome analysis of over 100 additional clones selected for carrying primary LOH tracts revealed that they too contained unselected structural alterations more often than control clones obtained without any selection. We also measured the rates of coincident LOH at two different chromosomes and found that double LOH formed at rates 14-150 fold higher than expected if the two underlying single LOH events occurred independently of each other. These results were consistent across different strain backgrounds, and in mutants incapable of entering meiosis. Our results indicate that a subset of mitotic cells within a population can experience discrete episodes of systemic genomic instability, when the entire genome becomes vulnerable and multiple chromosomal alterations can form over a narrow time window. They are reminiscent of early reports from the classic yeast genetics literature, as well as recent studies in humans, both in the cancer and genomic disorder contexts. The experimental model we describe provides a system to further dissect the fundamental biological processes responsible for punctuated bursts of structural genomic variation. SIGNIFICANCE STATEMENTMutations are generally thought to accumulate independently and gradually over many generations. Here, we combined complementary experimental approaches in budding yeast to track the appearance of chromosomal changes resulting in loss-of-heterozygosity (LOH). In contrast to the prevailing model, our results provide evidence for the existence of a path for non-independent accumulation of multiple chromosomal alteration events over few generations. These results are analogous to recent reports of bursts of genomic instability in human cells. The experimental model we describe provides a system to further dissect the fundamental biological processes underlying such punctuated bursts of mutation accumulation.

genetics↗

Genome evolution and pathoadaptation of Shigella

Shigella are pathogens originating within the Escherichia lineage but frequently classified as a separate genus. Shigella genomes contain numerous insertion sequences (ISs) that lead to pseudogenization of affected genes and an increase of non-homologous recombination. Here, we study 414 genomes of E. coli and Shigella strains to assess the contribution of genomic rearrangements to Shigella evolution. We found that Shigella experienced exceptionally high rates of intragenomic rearrangements and had a decreased rate of homologous recombination compared to pathogenic and non-pathogenic E. coli. The high rearrangement rate resulted in independent disruption of syntenic regions and parallel rearrangements in different Shigella lineages. Specifically, we identified two types of chromosomally encoded E3 ubiquitin-protein ligases acquired independently by all Shigella strains that also showed a high level of sequence conservation in the promoter and further in the 5 intergenic region. In the only available enteroinvasive E. coli (EIEC) strain, which is a pathogenic E. coli with a phenotype intermediate between Shigella and non-pathogenic E. coli, we found a rate of genome rearrangements comparable to those in other E. coli and no functional copies of the two Shigella-specific E3 ubiquitin ligases. These data indicate that accumulation of ISs influenced many aspects of genome evolution and played an important role in the evolution of intracellular pathogens. Our research demonstrates the power of comparative genomics-based on synteny block composition and an important role of non-coding regions in the evolution of genomic islands. ImportancePathogenic Escherichia coli strains frequently cause infections in humans. Many E. coli exist in nature and their ability to cause disease is fueled by their ability to incorporate novel genetic information by extensive horizontal gene transfer of plasmids and pathogenicity islands. The emergence of antibiotic-resistant Shigella spp., which are pathogenic forms of E. coli, coupled with the absence of an effective vaccine against them, highlights the importance of the continuing study of these pathogenic bacteria. Our study contributes to the understanding of genomic properties associated with molecular mechanisms underpinning the pathogenic nature of Shigella. We characterize the contribution of insertion sequences to the genome evolution of these intracellular pathogens and suggest a role of upstream regions of chromosomal ipaH genes in the Shigella pathogenesis. The methods of rearrangement analysis developed here are broadly applicable to the analysis of genotype-phenotype correlation in historically recently emerging bacterial pathogens.

bioinformatics↗

Surprising amount of stasis in repetitive genome content across the Brassicales

Genome size of plants has long piqued the interest of researchers due to the vast differences among organisms. However, the mechanisms that drive size differences have yet to be fully understood. Two important contributing factors to genome size are expansions of repetitive elements, such as transposable elements (TEs), and whole-genome duplications (WGD). Although studies have found correlations between genome size and both TE abundance and polyploidy, these studies typically test for these patterns within a genus or species. The plant order Brassicales provides an excellent system to test if genome size evolution patterns are consistent across larger time scales, as there are numerous WGDs. This order is also home to one of the smallest plant genomes, Arabidopsis thaliana - chosen as the model plant system for this reason - as well as to species with very large genomes. With new methods that allow for TE characterization from low-coverage genome shotgun data and 71 taxa across the Brassicales, we find no correlation between genome size and TE content, and more surprisingly we identify no significant changes to TE landscape following WGD.

evolutionary biology↗

The genome assembly and annotation of Magnolia biondii Pamp., a phylogenetically, economically, and medicinally important ornamental tree species

Magnolia biondii Pamp. (Magnoliaceae, magnoliids) is a phylogenetically, economically, and medicinally important ornamental tree species widely grown and cultivated in the north-temperate regions of China. Contributing a genome sequence for M. biondii will help resolve phylogenetic uncertainty of magnoliids and further understand individual trait evolution in Magnolia. We assembled a chromosome-level reference genome of M. biondii using ~67, ~175, and ~154 Gb of raw DNA sequences generated by Pacific Biosciences Single-molecule Real-time sequencing, 10X genomics Chromium, and Hi-C scaffolding strategies, respectively. The final genome assembly was 2.22 Gb with a contig N50 of 269.11 Kb and a BUSCO complete gene ratio of 91.90%. About 89.17% of the genome length was organized to 19 chromosomes, resulting in a scaffold N50 of 92.86 Mb. The genome contained 48,319 protein-coding genes, accounting for 22.97% of the genome length, in contrast to 66.48% of the genome length for the repetitive elements. We confirmed a Magnoliaceae specific WGD event that might have probably occurred shortly after the split of Magnoliaceae and Annonaceae. Functional enrichment of the Magnolia specific and expanded gene families highlighted genes involved in biosynthesis of secondary metabolites, plant-pathogen interaction, and response to stimulus, which may improve ecological fitness and biological adaptability of the lineage. Phylogenomic analyses recovered a sister relationship of magnoliids and Chloranthaceae, which are sister to a clade comprising monocots and eudicots. The genome sequence of M. biondii could empower trait improvement, germplasm conservation, and evolutionary studies on rapid radiation of early angiosperms.

plant biology↗

Genomic diversity generated by a transposable element burst in a rice recombinant inbred population

Genomes of all characterized higher eukaryotes harbor examples of transposable element (TE) bursts - the rapid amplification of TE copies throughout a genome. Despite their prevalence, understanding how bursts diversify genomes requires the characterization of actively transposing TEs before insertion sites and structural rearrangements have been obscured by selection acting over evolutionary time. In this study rice recombinant inbred lines (RILs), generated by crossing a bursting accession and the reference Nipponbare accession were exploited to characterize the spread of the very active Ping/mPing family through a small population and the resulting impact on genome diversity. Comparative sequence analysis of 272 individuals led to the identification of over 14,000 new insertions of the mPing miniature inverted-repeat transposable element (MITE) with no evidence for silencing of the transposase-encoding Ping element. In addition to new insertions, Ping-encoded transposase was found to preferentially catalyze the excision of mPing loci tightly linked to a second mPing insertion. Similarly, structural variations, including deletion of rice exons or regulatory regions, were enriched for those with breakpoints at one or both ends of linked mPing elements. Taken together, these results indicate that structural variations are generated during a TE burst as transposase catalyzes both the high copy numbers needed to distribute linked elements throughout the genome and the DNA cuts at the TE ends known to dramatically increase the frequency of recombination. Significance StatementTransposable elements (TEs) represent the largest component of the genomes of higher eukaryotes. Among this component are some TEs that have attained very high copy numbers with hundreds, even thousands of elements. By documenting the spread of mPing elements throughout the genomes of a rice population we demonstrate that such bursts of amplification generate functionally relevant genomic variations upon which selection can act. Specifically, continued mPing amplification increases the number of tightly linked elements that, in turn, increases the frequency of structural variations that appear to be derived from aberrant transposition events. The significance of this finding is that it provides a TE-mediated mechanism that may generate much of the structural variation represented by pan-genomes in plants and other organisms.

genetics↗

The transcriptional and splicing changes caused by hybridization can be globally recovered by genome doubling during allopolyploidization

Allopolyploidization, which involves hybridization and genome doubling, is a key driving force in higher plant evolution. The transcriptome reprogramming that accompanies allopolyploidization can cause extensive phenotypic variations, and thus confers allopolyploids higher evolutionary potential than their diploid progenitors. Despite many studies, little is known about the interplay between hybridization and genome doubling in transcriptome reprogramming during allopolyploidization. Here, we performed genome-wide analyses of gene expression and splicing changes during allopolyploidization in wheat and brassica lineages. Our results indicated that both hybridization and genome doubling can induce genome-wide transcriptional and splicing changes. Notably, the gene transcriptional and splicing changes caused by hybridization can be largely recovered to parental levels by genome doubling in allopolyploids. Since transcriptome reprogramming is an important contributor to heterosis, our results revealed that only part of the heterosis in hybrids can be fixed in allopolyploids through genome doubling. Therefore, our findings update the current understanding of the permanent fixation of heterosis in hybrids through genome doubling. In addition, our results indicated that a large proportion of the transcriptome reprogramming in interspecific hybrids was not caused by the merging of two parental genomes, providing novel insights into the mechanism of heterosis.

evolutionary biology↗

New evidence concerning the genome designations of the AC(DC) tetraploid Avena species

The tetraploid Avena species in the section Pachycarpa Baum, including A. insularis, A. maroccana, and A. murphyi, are thought to be involved in the evolution of hexaploid oats; however, their genome designations are still being debated. Repetitive DNA sequences play an important role in genome structuring and evolution, so understanding the chromosomal organization and distribution of these sequences in Avena species could provide valuable information concerning genome evolution in this genus. In this study, the chromosomal organizations and distributions of six repetitive DNA sequences (including three SSR motifs (TTC, AAC, CAG), one 5S rRNA gene fragment, and two oat A and C genome specific repeats) were investigated using non-denaturing fluorescence in situ hybridization (ND-FISH) in the three tetraploid species mentioned above and in two hexaploid oat species. Preferential distribution of the SSRs in centromeric regions was seen in the A and D genomes, whereas few signals were detected in the C genomes. Some intergenomic translocations were observed in the tetraploids; such translocations were also detected between the C and D genomes in the hexaploids. These results provide robust evidence for the presence of the D genome in all three tetraploids, strongly suggesting that the genomic constitution of these species is DC and not AC, as had been thought previously.

genetics↗

Broken, silent, and in hiding: Tamed endogenous pararetroviruses escape elimination from the genome of sugar beet (Beta vulgaris)

Background and AimsEndogenous pararetroviruses (EPRVs) are widespread components of plant genomes that originated from episomal DNA viruses of the Caulimoviridae family. Due to fragmentation and rearrangements, most EPRVs have lost their ability to replicate through reverse transcription and to initiate viral infection. Similar to the closely related retrotransposons, extant EPRVs were retained and often amplified in plant genomes for several million years. Here, we characterize the complete genomic EPRV fraction of the crop sugar beet (Beta vulgaris, Amaranthaceae) to understand how they shaped the beet genome and to suggest explanations for their absent virulence. MethodsUsing next- and third-generation sequencing data and the genome assembly, we reconstructed full-length in silico representatives for the three host-specific EPRV families (beetEPRVs) in the B. vulgaris genome. Focusing on the canonical family beetEPRV3, we investigated its chromosomal localization, abundance, and distribution by fluorescent in situ and Southern hybridization. Key ResultsBeetEPRVs range between 7.5 and 10.7 kb (0.3 % of the B. vulgaris genome) and are heterogeneous in structure and sequence. Although all three beetEPRV families were assigned to the florendoviruses, they showed variably arranged protein-coding domains, different degrees of fragmentation, and preferences for diverse sequence contexts. We observed small RNAs that target beetEPRVs in a family-specific manner, indicating stringent epigenetic suppression. We localized beetEPRV3 on all 18 sugar beet chromosomes, occurring preferentially in clusters and associated with heterochromatic, centromeric and intercalary satellite DNAs. BeetEPRV3 variants also exist in the genomes of related wild species, indicating an initial beetEPRV3 integration 13.4 to 7.2 million years ago. ConclusionsOur study in beet illustrates the variability of EPRV structure and sequence in a single host genome. Evidence of sequence fragmentation and epigenetic silencing imply possible plant strategies to cope with long-term persistence of EPRVs, including amplification, fixation in the heterochromatin, and containment of EPRV virulence.

plant biology↗

Sources of genomic diversity in the self-fertile plant pathogen, Sclerotinia sclerotiorum, and consequences for resistance breeding

The ascomycete, Sclerotinia sclerotiorum, has a broad host range and causes yield loss in dicotyledonous crops world wide. Genomic diversity and aggressiveness were determined in a population of 127 isolates from individual canola (Brassica napus) fields in western Canada. Genotyping with 39 simple sequence repeat (SSR) markers revealed each isolate was an unique haplotype. Analysis of molecular variation showed 97% was due to isolate and 3% to geographical location. Testing of mycelium compatibility identified clones of mutually compatible isolates, and stings of pairwise compatible isolates not seen before. Importantly, mutually compatible isolates had similar SSR haplotype, in contrast to high diversity among incompatible isolates. Isolates from the Province of Manitoba had higher allelic richness and higher mycelium compatibility (61%) than Alberta (35%) and Saskatchewan (39%). All compatible Manitoba isolates were interconnected in clones and strings, which can be explained by wetter growing seasons and more susceptible crops species both favouring more mycelium interaction and life cycles. Analysis of linkage disequilibrium rejected random recombination, consistent with a self-fertile fungus and restricted outcrossing due to mycelium incompatibility, and only one meiosis per lifecycle. More probable sources of genomic diversity is slippage during DNA replication and point mutation affecting single nucleotides, not withstanding the high mutation rate of SSRs compared to genes. It seems accumulation of these polymorphisms lead to increasing mycelium incompatibility in a population over time. A phylogenetic tree grouped isolates into 17 sub-populations. Aggressiveness was tested by inoculating one isolate from each sub-population onto B. napus lines with quantitative resistance. Results were significant for isolate, line, and isolate by line interaction. These isolates represent the genomic and pathogenic diversity in western Canada, and are suitable for resistance screening in canola breeding programs. Since the S. sclerotiorum life cycle is universal, conclusions on sources of genomic diversity extrapolates to populations in other geographical areas and host crops. Author summarySclerotinia sclerotiorum populations from various plant species and geographical areas have been studied extensively using mycelium compatibility tests and genotyping with a shared set of 6-13 SSR markers published in 2001. Most conclude the pathogen is clonally propagated with some degree of outcrossing. In the present study, a population of S. sclerotiorum isolates from 1.5 million km2 area in western Canada were tested for mycelium compatibility, and genotyped with 9 published and 30 newly developed SSR markers targeting all chromosomes in the dikaryot genome (8+8). A new way of visualizing mycelium compatibility results revealed clones of mutual compatible isolates, as well as long and short strings of pairwise compatible isolates. Importantly, clonal isolates had similar SSR haplotype, while incompatible isolates were highly dissimilar; a relationship difficult to discern previously. Analysis of population structure found a lack of linkage disequilibrium ruling out random recombination. Outcrossing, a result of alignment of non-sister chromosomes during meiosis, is unlikely in S. sclerotiorum, since mycelium incompatibility prevents karyogamy, and compatibility only occur between isolates with similar genomic composition. Instead, genomic diversity comprise transfer of nuclei through hyphal anastomosis, allelic modifications during cell division and point mutation. Genomic polymorphisms accumulate over time likely result in gradual divergence of individuals, which seems to resemble the ring-species concept. We are currently studying whether nuclei in microconidia might also contribute to diversity. A phylogenetic analysis grouped isolates into 17 sub-populations. One isolate from each sub-population showed different level of aggressiveness when inoculated onto B. napus lines previously determined to have quantitative resistance to a single isolate. Seed of these lines and S. sclerotiorum isolates have been transferred to plant breeders, and can be requested from the corresponding author for breeding purposes. Quantitative resistance is likely to hold up over time, since the rate of genomic change is relatively slow in S. sclerotiorum.

microbiology↗

Genome-scale sequencing and analysis of human, wolf and bison DNA from 25,000 year-old sediment

Archaeological sediments have been shown to preserve ancient DNA, but so far have not yielded genome-scale information of the magnitude of skeletal remains. We retrieved and analysed human and mammalian low-coverage nuclear and high-coverage mitochondrial genomes from Upper Palaeolithic sediments from Satsurblia cave, western Georgia, dated to 25,000 years ago. First, a human female genome with substantial basal Eurasian ancestry, which was an ancestry component of the majority of post-Ice Age people in the Near East, North Africa, and parts of Europe. Second, a wolf genome that is basal to extant Eurasian wolves and dogs and represents a previously unknown, likely extinct, Caucasian lineage that diverged from the ancestors of modern wolves and dogs before these diversified. Third, a bison genome that is basal to present-day populations, suggesting that population structure has been substantially reshaped since the Last Glacial Maximum. Our results provide new insights into the late Pleistocene genetic histories of these three species, and demonstrate that sediment DNA can be used not only for species identification, but also be a source of genome-wide ancestry information and genetic history. HighlightsO_LIWe demonstrate for the first time that genome sequencing from sediments is comparable to that of skeletal remains C_LIO_LIA single Pleistocene sediment sample from the Caucasus yielded three low-coverage mammalian ancient genomes C_LIO_LIWe show that sediment ancient DNA can reveal important aspects of the human and faunal past C_LIO_LIEvidence of an uncharacterized human lineage from the Caucasus before the Last Glacial Maximum C_LIO_LI[~]0.01-fold coverage wolf and bison genomes are both basal to present-day diversity, suggesting reshaping of population structure in both species C_LI

developmental biology↗

Constructing smaller genome graphs via string compression

The size of a genome graph -- the space required to store the nodes, their labels and edges -- affects the efficiency of operations performed on it. For example, the time complexity to align a sequence to a graph without a graph index depends on the total number of characters in the node labels and the number of edges in the graph. The size of the graph also affects the size of the graph index that is used to speed up the alignment. This raises the need for approaches to construct space-efficient genome graphs. We point out similarities in the string encoding approaches of genome graphs and the external pointer macro (EPM) compression model. Supported by these similarities, we present a pair of linear-time algorithms that transform between genome graphs and EPM-compressed forms. We show that the algorithms result in an upper bound on the size of the genome graph constructed based on an optimal EPM compression. In addition to the transformation, we show that equivalent choices made by EPM compression algorithms may result in different sizes of genome graphs. To further optimize the size of the genome graph, we purpose the source assignment problem that optimizes over the equivalent choices during compression and introduce an ILP formulation that solves that problem optimally. As a proof-of-concept, we introduce RLZ-Graph, a genome graph constructed based on the relative Lempel-Ziv EPM compression algorithm. We show that using RLZ-Graph, across all human chromosomes, we are able to reduce the disk space to store a genome graph on average by 40.7% compared to colored de Bruijn graphs constructed by Bifrost under the default settings. The RLZ-Graph software is available at https://github.com/Kingsford-Group/rlzgraph

bioinformatics↗

Rapid genomic convergent evolution in experimental populations of Trinidadian guppies (Poecilia reticulata)

It is now accepted that phenotypic evolution can occur quickly but the genetic basis of rapid adaptation to natural environments is largely unknown in multicellular organisms. Population genomic studies of experimental populations of Trinidadian guppies (Poecilia reticulata) provide a unique opportunity to study this phenomenon. Guppy populations that were transplanted from high-predation (HP) to low-predation (LP) environments have been shown to mimic naturally-colonised LP populations phenotypically in as few as 8 generations. The new phenotypes persist in subsequent generations in lab environments, indicating their high heritability. Here, we compared whole genome variation in four populations recently introduced into LP sites along with the corresponding HP source population. We examined genome-wide patterns of genetic variation to estimate past demography, and uncovered signatures of selection with a combination of genome scans and a novel multivariate approach based on allele frequency change vectors. We were able to identify a limited number of candidate loci for convergent evolution across the genome. In particular, we found a region on chromosome 15 under strong selection in three of the four populations, with our multivariate approach revealing subtle parallel changes in allele frequency in all four populations across this region. Investigating patterns of genome-wide selection in this uniquely replicated experiment offers remarkable insight into the mechanisms underlying rapid adaptation, providing a basis for comparison with other species and populations experiencing rapidly changing environments. IMPACT STATEMENTThe genetic basis of rapid adaptation to new environments is largely unknown. Here we take advantage of a unique replicated experiment in the wild, where guppies from a high predation source were introduced into four low predation localities. Previous reports document census size fluctuations and rapid phenotypic evolution in these populations. We used genome-wide sequencing to understand past demography and selection. We detected clear signals of population growth and bottlenecks at the genome-wide level matching known census population data changes. We then identified candidate regions of selection across the genome, some of which were shared between populations. In particular, using a novel multivariate method, we identified parallel allele frequency change at a strong candidate locus for adaptation to low predation. These results and methods will be of use to those studying evolution at a recent, ecological timescale.

evolutionary biology↗

Comparative repeat profiling of two closely related conifers (Larix decidua and Larix kaempferi) reveals high genome similarity with only one fast-evolving satellite DNA

In eukaryotic genomes, cycles of repeat expansion and removal lead to large-scale genomic changes and propel organisms forward in evolution. However, in conifers, active repeat removal is thought to be limited, leading to expansions of their genomes, mostly exceeding 10 gigabasepairs. As a result, conifer genomes are largely littered with fragmented and decayed repeats. Here, we aim to investigate how the repeat landscapes of two related conifers have diverged, given the conifers accumulative genome evolution mode. For this, we applied low coverage sequencing and read clustering to the genomes of European and Japanese larch, Larix decidua (Lamb.) Carriere and Larix kaempferi (Mill.), that arose from a common ancestor, but are now geographically isolated. We found that both Larix species harbored largely similar repeat landscapes, especially regarding the transposable element content. To pin down possible genomic changes, we focused on the repeat class with the fastest sequence turnover: satellite DNAs (satDNAs). Using comparative bioinformatics, Southern, and fluorescent in situ hybridization, we reveal the satDNAs organizational patterns, their abundances, and chromosomal locations. Four out of the five identified satDNAs are widespread in the Larix genus, with two even present in the more distantly related Pseudotsuga and Abies genera. Unexpectedly, the EulaSat3 family was restricted to L. decidua and absent from L. kaempferi, indicating its evolutionarily young age. Taken together, our results exemplify how the accumulative genome evolution of conifers may limit the overall divergence of repeats after speciation, producing only few repeat-induced genomic novelties.

plant biology↗