bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Bacterial genes outnumber archaeal genes in eukaryotic genomes

The origin of eukaryotes is one of evolutions most important transitions, yet it is still poorly understood. Evidence for how it occurred should be preserved in eukaryotic genomes. Based on phylogenetic trees from ribosomal RNA and ribosomal proteins, eukaryotes are typically depicted as branching together with or within archaea. This ribosomal affiliation is widely interpreted as evidence for an archaeal origin of eukaryotes. However, the extent to which the archaeal ancestry of genes for the cytosolic ribosomes of eukaryotic cells is representative for the rest of the eukaryotic genome is unknown. Here we have clustered 19,050,992 protein sequences from 5,443 bacteria and 212 archaea with 3,420,731 protein sequences from 150 eukaryotes spanning six eukaryotic supergroups to identify genes that link eukaryotes exclusively to bacteria and archaea respectively. By downsampling the bacterial sample we obtain estimates for the bacterial and archaeal proportions of genes among 150 eukaryotic genomes. Eukaryotic genomes possess a bacterial majority of genes. On average, eukaryotic genes are 56% bacterial in origin. The majority drops to 53% in eukaryotes that never possessed plastids, and increases to 61% in photosynthetic eukaryotic lineages, where the cyanobacterial ancestor of plastids contributed additional genes to the eukaryotic genome, reaching 67% in higher plants. Intracellular parasites, which undergo reductive evolution in adaptation to the nutrient rich environment of the cells that they infect, relinquish bacterial genes for metabolic processes. In the current sample, this process of adaptive gene loss is most pronounced in the human parasite Encephalitozoon intestinalis with 86% archaeal and 14% bacterial derived genes. The most bacterial eukaryote genome sampled is rice, with 67% bacterial and 33% archaeal genes. The functional dichotomy, initially described for yeast, of archaeal genes being involved in genetic information processing and bacterial genes being involved in metabolic processes is conserved across all eukaryotic supergroups.

evolutionary biology

Horizontal gene transfer to a defensive symbiont with a reduced genome amongst a multipartite beetle microbiome

The loss of functions required for independent life when living within a host gives rise to reduced genomes in obligate bacterial symbionts. Although this phenomenon can be explained by existing evolutionary models, its initiation is not well understood. Here, we describe the microbiome associated with eggs of the beetle Lagria villosa, containing multiple bacterial symbionts related to Burkholderia gladioli including a reduced-genome symbiont thought to produce the defensive compound lagriamide. We find that the putative lagriamide producer is the only symbiont undergoing genome reduction, and that it has already lost most primary metabolism and DNA repair pathways. The horizontal acquisition of the lagriamide biosynthetic gene cluster likely preceded genome reduction, and unexpectedly we found that the symbiont accepted additional genes horizontally during genome reduction, even though it lacks the capacity for homologous recombination. These horizontal gene transfers suggest that absolute genetic isolation is not a requirement for genome reduction.

evolutionary biology

EAGLE: an algorithm that utilizes a small number of genomic features to predict tissue/cell type-specific enhancer-gene interaction

Long-range regulation by distal enhancers is crucial for many biological processes. The existing methods for enhancer-target gene prediction often require many genomic features. This makes them difficult to be applied to many cell types, in which the relevant datasets are not always available. Here, we design a tool EAGLE, an enhancer and gene learning ensemble method for identification of Enhancer-Gene (EG) interactions. Unlike existing tools, EAGLE used only six features derived from the genomic features of enhancers and gene expression datasets. Cross-validation revealed that EAGLE outperformed other existing methods. Enrichment analyses on special transcriptional factors, epigenetic modifications, and eQTLs demonstrated that EAGLE could distinguish the interacting pairs from non- interacting ones. Finally, EAGLE was applied to mouse and human genomes and identified 7,680,203 and 7,437,255 EG interactions involving 31,375 and 43,724 genes, 138,547 and 177,062 enhancers across 89 and 110 tissue/cell types in mouse and human, respectively. The obtained interactions are accessible through an interactive database enhanceratlas.org. The EAGLE method is available at https://github.com/EvansGao/EAGLE and the predicted datasets are available in http://www.enhanceratlas.org/.\n\nAuthor summaryEnhancers are DNA sequences that interact with promoters and activate target genes. Since enhancers often located far from the target genes and the nearest genes are not always the targets of the enhancers, the prediction of enhancer-target gene relationships is a big challenge. Although a few computational tools are designed for the prediction of enhancer-target genes, its difficult to apply them in most tissue/cell types due to a lack of enough genomic datasets. Here we proposed a new method, EAGLE, which utilizes a small number of genomic features to predict tissue/cell type-specific enhancer-gene interactions. Comparing with other existing tools, EAGLE displayed a better performance in the 10-fold cross-validation and cross-sample test. Moreover, the predictions by EAGLE were validated by other independent evidence such as the enrichment of relevant transcriptional factors, epigenetic modifications, and eQTLs.\n\nFinally, we integrated the enhancer-target relationships obtained from human and mouse genomes into an interactive database EnhancerAtlas, http://www.enhanceratlas.org/.

bioinformatics

TypeTE: a tool to genotype mobile element insertions from whole genome resequencing data

Alu retrotransposons account for more than 10% of the human genome, and insertions of these elements create structural variants segregating in human populations. Such polymorphic Alu are powerful markers to understand population structure, and they represent variants that can greatly impact genome function, including gene expression. Accurate genotyping of Alu and other mobile elements has been challenging. Indeed, we found that Alu genotypes previously called for the 1000 Genomes Project are sometimes erroneous, which poses significant problems for phasing these insertions with other variants that comprise the haplotype. To ameliorate this issue, we introduce a new pipeline -- TypeTE -- which genotypes Alu insertions from whole-genome sequencing data. Starting from a list of polymorphic Alus, TypeTE identifies the hallmarks (poly-A tail and target site duplication) and orientation of Alu insertions using local re-assembly to reconstruct presence and absence alleles. Genotype likelihoods are then computed after re-mapping sequencing reads to the reconstructed alleles. Using a gold standard set of PCR-based genotyping of >200 loci, we show that TypeTE improves genotype accuracy from 83% to 92% in the 1000 Genomes dataset. TypeTE can be readily adapted to other retrotransposon families and brings a valuable toolbox addition for population genomics.

bioinformatics

Big data reveals fewer recombination hotspots than expected in human genome

Recombination is a major force that shapes genetic diversity. Determination of recombination rate is important and can theoretically be improved by increasing the sample size. However, it is challenging to estimate recombination rates when the sample size is extraordinarily large because of computational burden. In this study, we used a refined artificial intelligence approach to estimate the recombination rate of the human genome using the UK10K human genomic dataset with 7,562 genomic sequences and its three subsets with 200, 400 and 2,000 genomic sequences under the Out-of-Africa demography model. We not only obtained an accurate human genetic map, but also found that the fluctuation of estimated recombination rate is reduced along the human genome when the sample size is increased. UK10K recombination activity is less concentrated than its subsets. Our results demonstrate how the sample size affects the estimated recombination rate, and analyses of a larger number of genomes result in a more precise estimation of recombination rate.

genetics

Importance of parental genome balance in the generation of novel yet heritable epigenetic and transcriptional states during doubled haploid breeding

BackgroundDoubling the genome contribution of haploid plants has accelerated breeding in most cultivated crop species. Although plant doubled haploids are isogenic in nature, they frequently display unpredictable phenotypes, thus limiting the potential of this technology. Therefore, being able to predict the factors implicated in this phenotypic variability could accelerate the generation of desirable genomic combinations and ultimately plant breeding.\n\nResultsWe use computational analysis to assess the transcriptional and epigenetic dynamics taking place during doubled haploids generation in the genome of Brassica oleracea. We observe that doubled haploid lines display unexpected levels of transcriptional and epigenetic variation, and that this variation is largely due to imbalanced contribution of parental genomes. We reveal that epigenetic modification of transposon-related sequences during DH breeding contributes to the generation of unpredictable yet heritable transcriptional states. Targeted epigenetic manipulation of these elements using dCas9-hsTET3 confirms their role in transcriptional regulation. We have uncovered a hitherto unknown role for parental genome balance in the transcriptional and epigenetic stability of doubled haploids.\n\nConclusionsThis is the first study that demonstrates the importance of parental genome balance in the transcriptional and epigenetic stability of doubled haploids, thus enabling predictive models to improve doubled haploid-assisted plant breeding.

plant biology

Accounting for diverse evolutionary forces reveals the mosaic nature of selection on genomic regions associated with human preterm birth

Human pregnancy requires the coordinated function of multiple tissues in both mother and fetus and has evolved in concert with major human adaptations. As a result, pregnancy-associated phenotypes and related disorders are genetically complex and have likely been sculpted by diverse evolutionary forces. However, there is no framework to comprehensively evaluate how these traits evolved or to explore the relationship of evolutionary signatures on trait-associated genetic variants to molecular function. Here we develop an approach to test for signatures of diverse evolutionary forces, including multiple types of selection, and apply it to genomic regions associated with spontaneous preterm birth (sPTB), a complex disorder of global health concern. We find that sPTB-associated regions harbor diverse evolutionary signatures including evolutionary sequence conservation (consistent with the action of negative selection), excess population differentiation (local adaptation), accelerated evolution (positive selection), and balanced polymorphism (balancing selection). Furthermore, these genomic regions show diverse functional characteristics which enables us to use evolutionary and molecular lines of evidence to develop hypotheses about how these genomic regions contribute to sPTB risk. In summary, we introduce an approach for inferring the spectrum of evolutionary forces acting on genomic regions associated with complex disorders. When applied to sPTB-associated genomic regions, this approach both improves our understanding of the potential roles of these regions in pathology and illuminates the mosaic nature of evolutionary forces acting on genomic regions associated with sPTB.

evolutionary biology

Finding genetic variants in plants without complete genomes

Structural variants and presence/absence polymorphisms are common in plant genomes, yet they are routinely overlooked in genome-wide association studies (GWAS). Here, we expand the genetic variants detected in GWAS to include major deletions, insertions, and rearrangements. We first use raw sequencing data directly to derive short sequences, k-mers, that mark a broad range of polymorphisms independently of a reference genome. We then link k-mers associated with phenotypes to specific genomic regions. Using this approach, we re-analyzed 2,000 traits measured in Arabidopsis thaliana, tomato, and maize populations. Associations identified with k-mers recapitulate those found with single-nucleotide polymorphisms (SNPs), however, with stronger statistical support. Moreover, we identified new associations with structural variants and with regions missing from reference genomes. Our results demonstrate the power of performing GWAS before linking sequence reads to specific genomic regions, which allow detection of a wider range of genetic variants responsible for phenotypic variation.

plant biology

Machine Learning Reveals Spatiotemporal Genome Evolution in Asian Rice Domestication

Domestication is anthropogenic evolution that fulfills mankinds critical food demand. As such, elucidating the molecular mechanisms behind this process promotes the development of future new food resources including crops. With the aim of understanding the long-term domestication process of Asian rice and by employing the Oryza sativa subspecies (indica and japonica) as an Asian rice domestication model, we scrutinized past genomic introgressions between them as traces of domestication. Here we show the genome-wide introgressive region (IR) map of Asian rice, by utilizing 4,587 accession genotypes with a stable outgroup species, particularly at the finest resolution through a machine learning-aided method. The IR map revealed that 14.2% of the rice genome consists of IRs, including both wide IRs (recent) and narrow IRs (ancient). This introgressive landscape with their time calibration indicates that introgression events happened in multiple genomic regions over multiple periods. From the correspondence between our wide IRs and the so-called selective sweep regions, we provide a definitive answer to a long-standing controversy over the evolutionary origin of Asian rice domestication, single or multiple origins: It heavily depends upon which regions you pay attention to, implying that wider genomic regions represent immediate short history of Asian rice domestication as a likely support to the single origin, while its ancient history is interspersed in narrower traces throughout the genome as a possible support to the multiple origin.

plant biology

On the impact of contaminants on the accuracy of genome skimming and the effectiveness of exclusion read filters

The ability to detect the identity of a sample obtained from its environment is a cornerstone of molecular ecological research. Thanks to the falling price of shotgun sequencing, genome skimming, the acquisition of short reads spread across the genome at low coverage, is emerging as an alternative to traditional barcoding. By obtaining far more data across the whole genome, skimming has the promise to increase the precision of sample identification beyond traditional barcoding while keeping the costs manageable. While methods for assembly-free sample identification based on genome skims are now available, little is known about how these methods react to the presence of DNA from organisms other than the target species. In this paper, we show that the accuracy of distances computed between a pair of genome skims based on k-mer similarity can degrade dramatically if the skims include contaminant reads; i.e., any reads originating from other organisms. We establish a theoretical model of the impact of contamination. We then suggest and evaluate a solution to the contamination problem: Query reads in a genome skim against an extensive database of possible contaminants (e.g., all microbial organisms) and filter out any read that matches. We evaluate the effectiveness of this strategy when implemented using Kraken-II, in detailed analyses. Our results show substantial improvements in accuracy as a result of filtering but also point to limitations, including a need for relatively close matches in the contaminant database.

bioinformatics

BlobToolKit Interactive quality assessment of genome assemblies

Reconstruction of target genomes from sequence data produced by instruments that are agnostic as to the species-of-origin may be confounded by contaminant DNA. Whether introduced during sample processing or through co-extraction alongside the target DNA, if insufficient care is taken during the assembly process, the final assembled genome may be a mixture of data from several species. Such assemblies can confound sequence-based biological inference and, when deposited in public databases, may be included in downstream analyses by users unaware of underlying problems. We present BlobToolKit, a software suite to aid researchers in identifying and isolating non-target data in draft and publicly available genome assemblies. BlobToolKit can be used to process assembly, read and analysis files for fully reproducible interactive exploration in the browser-based Viewer. BlobToolKit can be used during assembly to filter non-target DNA, helping researchers produce assemblies with high biological credibility. We have been running an automated BlobToolKit pipeline on eukaryotic assemblies publicly available in the International Nucleotide Sequence Data Collaboration and are making the results available through a public instance of the Viewer at https://blobtoolkit.genomehubs.org/view. We aim to complete analysis of all publicly available genomes and then maintain currency with the flow of new genomes. We have worked to embed these views into the presentation of genome assemblies at the European Nucleotide Archive, providing an indication of assembly quality alongside the public record with links out to allow full exploration in the Viewer.

bioinformatics

mity: A highly sensitive mitochondrial variant analysis pipeline for whole genome sequencing data

MotivationMitochondrial diseases (MDs) are the most common group of inherited metabolic disorders and are often challenging to diagnose due to extensive genotype-phenotype heterogeneity. MDs are caused by mutations in the nuclear or mitochondrial genome, where pathogenic mitochondrial variants are usually heteroplasmic and typically at much lower allelic fraction in the blood than affected tissues. Both genomes can now be readily analysed using unbiased whole genome sequencing (WGS), but most nuclear variant detection methods fail to detect low heteroplasmy variants in the mitochondrial genome. ResultsWe present mity, a bioinformatics pipeline for detecting and interpreting heteroplasmic SNVs and INDELs in the mitochondrial genome using WGS data. In 2,980 healthy controls, we observed on average 3,166x coverage in the mitochondrial genome using WGS from blood. mity utilises this high depth to detect pathogenic mitochondrial variants, even at low heteroplasmy. mity enables easy interpretation of mitochondrial variants and can be incorporated into existing diagnostic WGS pipelines. This could simplify the diagnostic pathway, avoid invasive tissue biopsies and increase the diagnostic rate for MDs and other conditions caused by impaired mitochondrial function. Availabilitymity is available from https://github.com/KCCG/mityunder an MIT license. Contactclare.puttick@crick.ac.uk, carolyn.sue@sydney.edu.au, MCowley@ccia.org.au

bioinformatics

Comparative and population genomics approaches reveal the basis of adaptation to deserts in a small rodent

Organisms that live in deserts offer the opportunity to investigate how species adapt to environmental conditions that are lethal to most plants and animals. In the hot deserts of North America, high temperatures and lack of water are conspicuous challenges for organisms living there. The cactus mouse (Peromyscus eremicus) displays several adaptations to these conditions, including low metabolic rate, heat tolerance, and the ability to maintain homeostasis under extreme dehydration. To investigate the genomic basis of desert adaptation in cactus mice, we built a chromosome-level genome assembly and resequenced 26 additional cactus mouse genomes from two locations in southern California (USA). Using these data, we integrated comparative, population, and functional genomic approaches. We identified 16 gene families exhibiting significant contractions or expansions in the cactus mouse compared to 17 other Myodontine rodent genomes, and found 232 sites across the genome associated with selective sweeps. Functional annotations of candidate gene families and selective sweeps revealed a pervasive signature of selection at genes involved in the synthesis and degradation of proteins, consistent with the evolution of cellular mechanisms to cope with protein denaturation caused by thermal and hyperosmotic stress. Other strong candidate genes included receptors for bitter taste, suggesting a dietary shift towards chemically defended desert plants and insects, and a growth factor involved in lipid metabolism, potentially involved in prevention of dehydration. Understanding how species adapted to the recent emergence of deserts in North America will provide an important foundation for predicting future evolutionary responses to increasing temperatures, droughts and desertification in the cactus mouse and other species.

evolutionary biology

wg-blimp: an end-to-end analysis pipeline for whole genome bisulfite sequencing data

BackgroundAnalysing whole genome bisulfite sequencing datasets is a data-intensive task that requires comprehensive and reproducible workflows to generate valid results. While many algorithms have been developed for tasks such as alignment, comprehensive end-to-end pipelines are still sparse. Furthermore, previous pipelines lack features or show technical deficiencies, thus impeding analyses. ResultsWe developed wg-blimp (whole genome bisulfite sequencing methylation analysis pipeline) as an end-to-end pipeline to ease whole genome bisulfite sequencing data analysis. It integrates established algorithms for alignment, quality control, methylation calling, detection of differentially methylated regions, and methylome segmentation, requiring only a reference genome and raw sequencing data as input. Comparing wg-blimp to previous end-to-end pipelines reveals similar setups for common sequence processing tasks, but shows differences for post-alignment analyses. We improve on previous pipelines by providing a more comprehensive analysis workflow as well as an interactive user interface. To demonstrate wg-blimps ability to produce correct results we used it to call differentially methylated regions for two publicly available datasets. We were able to replicate 112 of 114 previously published regions, and found results to be consistent with previous findings. We further applied wg-blimp to a publicly available sample of embryonic stem cells to showcase methylome segmentation. As expected, unmethylated regions were in close proximity of transcription start sites. Segmentation results were consistent with previous analyses, despite different reference genomes and sequencing techniques. Conclusionswg-blimp provides a comprehensive analysis pipeline for whole genome bisulfite sequencing data as well as a user interface for simplified result inspection. We demonstrated its applicability by analysing multiple publicly available datasets. Thus, wg-blimp is a relevant alternative to previous analysis pipelines and may facilitate future epigenetic research.

bioinformatics

Population genetics of the coral Acropora millepora: Towards a genomic predictor of bleaching

Although reef-building corals are rapidly declining worldwide, responses to bleaching vary both within and among species. Because these inter-individual differences are partly heritable, they should in principle be predictable from genomic data. Towards that goal, we generated a chromosome-scale genome assembly for the coral Acropora millepora. We then obtained whole genome sequences for 237 phenotyped samples collected at 12 reefs distributed along the Great Barrier Reef, among which we inferred very little population structure. Scanning the genome for evidence of local adaptation, we detected signatures of long-term balancing selection in the heat-shock co-chaperone sacsin. We further used 213 of the samples to conduct a genome-wide association study of visual bleaching score, incorporating the polygenic score derived from it into a predictive model for bleaching in the wild. These results set the stage for the use of genomics-based approaches in conservation strategies.

evolutionary biology

N6-methyladenosine regulates Influenza A virus mRNA stability yet is rarely found on genomic RNA

Previous studies have found widespread N6-methyladenosine (m6A methylation) on all forms of Influenza A virus (IAV) RNA, with m6A found critical for viral replication, pathogenicity as well as viral RNA packaging. Here we applied the latest quantitative technologies to revisit the methylation landscape on the anti-sense genomic RNA of IAV. Unexpectedly, upon Ultra-Performance Liquid Chromatography-Tandem Mass Spectrometry (UPLC-MS/MS) analysis of IAV virion -extracted genomic RNA, we detected very little m6A regardless of production from human cells or chicken eggs. Concordantly, Nanopore direct RNA sequencing also detected an overall low occurrence and stoichiometry (generally <5%) of m6A across all viral genomic RNA segments, compared with abundant m6A sites on viral mRNAs at ~20-30% m6A. Cross validation with glyoxal- and nitrite-mediated deamination of unmethylated adenosines (GLORI) confirmed multiple m6A sites on viral mRNA yet very few m6A on the genomic RNA. This paucity of m6A on genomic RNA makes it unlikely that m6A contributes to viral RNA packaging. Knockdown or pharmacological inhibition of the m6A methyltransferase METTL3 as well as the reader protein YTHDF2 both reduced viral mRNA levels and infectious viral particle production, with YTHDF2 promoting viral mRNA stability. Thus, the presence of m6A on IAV transcripts is indeed proviral, yet it is the mRNAs instead of genomic RNAs that are methylated at functionally relevant levels. Lastly, we provide proof of concept that a METTL3 small molecule inhibitor can be antiviral, and propose that m6A-targeted antivirals would mainly impact the intracellular gene expression phase of IAV replication.

microbiology

Optical and physical mapping with local finishing enables megabase-scale resolution of agronomically important regions in the wheat genome

BackgroundNumerous scaffold-level sequences for wheat are now being released and, in this context, we report on a strategy for improving the overall assembly to a level comparable to that of the human genome.\n\nResultsUsing chromosome 7A of wheat as a model, sequence-finished megabase scale sections of this chromosome were established by combining a new independent assembly based on a BAC-based physical map, BAC pool paired end sequencing, chromosome arm specific mate-pair sequencing and Bionano optical mapping with the IWGSC RefSeq v1.0 sequence and its underlying raw data. The combined assembly results in 18 super-scaffolds across the chromosome. The value of finished genome regions is demonstrated for two approximately 2.5 Mb regions associated with yield and the grain quality phenotype of fructan carbohydrate grain levels. In addition, the 50 Mb centromere region analysis incorporates cytological data highlighting the importance of non-sequence data in the assembly of this complex genome region.\n\nConclusionsSufficient genome sequence information is shown to be now available for the wheat community to produce sequence-finished releases of each chromosome of the reference genome. The high-level completion identified that an array of seven fructosyl transferase genes underpins grain quality and yield attributes are affected by five f-box-only-protein-ubiquitin ligase domain and four root-specific lipid transfer domain genes. The completed sequence also includes the centromere.

genomics

Complete microbial genomes for public health in Australia and Southwest Pacific

Complete genomes of microbial pathogens are essential for the phylogenomic analyses that increasingly underpin core public health lab activities. Here, we present complete genomes of pathogen strains of regional importance to the Southwest Pacific and Australia. These enrich the catalogue of globally available complete genomes for public health while providing valuable strains to regional public health labs.\n\nAnnouncementWhole-genome sequence (WGS) data is increasingly important in public health microbiology (1-4). The data can be used to replicate many of the basic bacterial sub-typing approaches, as well as support epidemiological investigations, such as surveillance and outbreak investigation (5-7). The appeal of WGS data comes from the promise of a single workflow to process all microbial pathogens that can provide easily portable data that promotes deeper integration of surveillance and investigation efforts across jurisdictions. This promise is leading to a concerted effort to move microbial public health to a primarily genome-based workflow at numerous jurisdictions (8-10), including Australia (11).

genomics