bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

N6-methyladenosine regulates Influenza A virus mRNA stability yet is rarely found on genomic RNA

Previous studies have found widespread N6-methyladenosine (m6A methylation) on all forms of Influenza A virus (IAV) RNA, with m6A found critical for viral replication, pathogenicity as well as viral RNA packaging. Here we applied the latest quantitative technologies to revisit the methylation landscape on the anti-sense genomic RNA of IAV. Unexpectedly, upon Ultra-Performance Liquid Chromatography-Tandem Mass Spectrometry (UPLC-MS/MS) analysis of IAV virion -extracted genomic RNA, we detected very little m6A regardless of production from human cells or chicken eggs. Concordantly, Nanopore direct RNA sequencing also detected an overall low occurrence and stoichiometry (generally <5%) of m6A across all viral genomic RNA segments, compared with abundant m6A sites on viral mRNAs at ~20-30% m6A. Cross validation with glyoxal- and nitrite-mediated deamination of unmethylated adenosines (GLORI) confirmed multiple m6A sites on viral mRNA yet very few m6A on the genomic RNA. This paucity of m6A on genomic RNA makes it unlikely that m6A contributes to viral RNA packaging. Knockdown or pharmacological inhibition of the m6A methyltransferase METTL3 as well as the reader protein YTHDF2 both reduced viral mRNA levels and infectious viral particle production, with YTHDF2 promoting viral mRNA stability. Thus, the presence of m6A on IAV transcripts is indeed proviral, yet it is the mRNAs instead of genomic RNAs that are methylated at functionally relevant levels. Lastly, we provide proof of concept that a METTL3 small molecule inhibitor can be antiviral, and propose that m6A-targeted antivirals would mainly impact the intracellular gene expression phase of IAV replication.

microbiology

Optical and physical mapping with local finishing enables megabase-scale resolution of agronomically important regions in the wheat genome

BackgroundNumerous scaffold-level sequences for wheat are now being released and, in this context, we report on a strategy for improving the overall assembly to a level comparable to that of the human genome.\n\nResultsUsing chromosome 7A of wheat as a model, sequence-finished megabase scale sections of this chromosome were established by combining a new independent assembly based on a BAC-based physical map, BAC pool paired end sequencing, chromosome arm specific mate-pair sequencing and Bionano optical mapping with the IWGSC RefSeq v1.0 sequence and its underlying raw data. The combined assembly results in 18 super-scaffolds across the chromosome. The value of finished genome regions is demonstrated for two approximately 2.5 Mb regions associated with yield and the grain quality phenotype of fructan carbohydrate grain levels. In addition, the 50 Mb centromere region analysis incorporates cytological data highlighting the importance of non-sequence data in the assembly of this complex genome region.\n\nConclusionsSufficient genome sequence information is shown to be now available for the wheat community to produce sequence-finished releases of each chromosome of the reference genome. The high-level completion identified that an array of seven fructosyl transferase genes underpins grain quality and yield attributes are affected by five f-box-only-protein-ubiquitin ligase domain and four root-specific lipid transfer domain genes. The completed sequence also includes the centromere.

genomics

Improved genome inference in the MHC using a population reference graph

In humans and many other species, while much is known about the extent and structure of genetic variation, such information is typically not used in assembling novel genomes. Rather, a single reference is used against which to map reads, which can lead to poor characterisation of regions of high sequence or structural diversity. Here, we introduce a population reference graph, which combines multiple reference sequences as well as catalogues of SNPs and short indels. The genomes of novel samples are reconstructed as paths through the graph using an efficient hidden Markov Model, allowing for recombination between different haplotypes and variants. By applying the method to the 4.5Mb extended MHC region on chromosome 6, combining eight assembled haplotypes, sequences of known classical HLA alleles and 87,640 SNP variants from the 1000 Genomes Project, we demonstrate, using simulations, SNP genotyping, short-read and longread data, how the method improves the accuracy of genome inference. Moreover, the analysis reveals regions where the current set of reference sequences is substantially incomplete, particularly within the Class II region, indicating the need for continued development of reference-quality genome sequences.

Genomics

Predicting genome sizes and restriction enzyme recognition-sequence probabilities across the eukaryotic tree of life

High-throughput sequencing of reduced representation libraries obtained through digestion with restriction enzymes - generically known as restriction-site associated DNA sequencing (RAD-seq) - is a common strategy to generate genome-wide genotypic and sequence data from eukaryotes. A critical design element of any RAD-seq study is a knowledge of the approximate number of genetic markers that can be obtained for a taxon using different restriction enzymes, as this number determines the scope of a project, and ultimately defines its success. This number can only be directly determined if a reference genome sequence is available, or it can be estimated if the genome size and restriction recognition sequence probabilities are known. However, both scenarios are uncommon for non-model species. Here, we performed systematic in silico surveys of recognition sequences, for diverse and commonly used type II restriction enzymes across the eukaryotic tree of life. Our observations reveal that recognition-sequence frequencies for a given restriction enzyme are strikingly variable among broad eukaryotic taxonomic groups, being largely determined by phylogenetic relatedness. We demonstrate that genome sizes can be predicted from cleavage frequency data obtained with restriction enzymes targeting neutral elements. Models based on genomic compositions are also effective tools to accurately calculate probabilities of recognition sequences across taxa, and can be applied to species for which reduced-representation data is available (including transcriptomes and neutral RAD-seq datasets). The analytical pipeline developed in this study, PredRAD (https://github.com/phrh/PredRAD), and the resulting databases constitute valuable resources that will help guide the design of any study using RAD-seq or related methods.

Genomics

A Comprehensive Assessment of Somatic Mutation Calling in Cancer Genomes

The emergence of next generation DNA sequencing technology is enabling high-resolution cancer genome analysis. Large-scale projects like the International Cancer Genome Consortium (ICGC) are systematically scanning cancer genomes to identify recurrent somatic mutations. Second generation DNA sequencing, however, is still an evolving technology and procedures, both experimental and analytical, are constantly changing. Thus the research community is still defining a set of best practices for cancer genome data analysis, with no single protocol emerging to fulfil this role. Here we describe an extensive benchmark exercise to identify and resolve issues of somatic mutation calling. Whole genome sequence datasets comprising tumor-normal pairs from two different types of cancer, chronic lymphocytic leukaemia and medulloblastoma, were shared within the ICGC and submissions of somatic mutation calls were compared to verified mutations and to each other. Varying strategies to call mutations, incomplete awareness of sources of artefacts, and even lack of agreement on what constitutes an artefact or real mutation manifested in widely varying mutation call rates and somewhat low concordance among submissions. We conclude that somatic mutation calling remains an unsolved problem. However, we have identified many issues that are easy to remedy that are presented here. Our study highlights critical issues that need to be addressed before this valuable technology can be routinely used to inform clinical decision-making.\n\nAbbreviations and Definitionsaligner = mapper, these terms are used interchangeably

Genomics

Mapping bias overestimates reference allele frequencies at the HLA genes in the 1000 Genomes Project phase I data

Next Generation Sequencing (NGS) technologies have become the standard for data generation in studies of population genomics, as the 1000 Genomes Project (1000G). However, these techniques are known to be problematic when applied to highly polymorphic genomic regions, such as the Human Leukocyte Antigen (HLA) genes. Because accurate genotype calls and allele frequency estimations are crucial to population ge-nomics analises, it is important to assess the reliability of NGS data. Here, we evaluate the reliability of genotype calls and allele frequency estimates of the SNPs reported by 1000G (phase I) at five HLA genes (HLA-A, -B, -C, -DRB1, -DQB1). We take advantage of the availability of HLA Sanger sequencing of 930 of the 1,092 1000G samples, and use this as a gold standard to benchmark the 1000G data. We document that 18.6% of SNP genotype calls in HLA genes are incorrect, and that allele frequencies are estimated with an error higher than {+/-}0.1 at approximately 25% of the SNPs in HLA genes. We found a bias towards overestimation of reference allele frequency for the 1000G data, indicating mapping bias is an important cause of error in frequency estimation in this dataset. We provide a list of sites that have poor allele frequency estimates, and discuss the outcomes of including those sites in different kinds of analyses. Since the HLA region is the most polymorphic in the human genome, our results provide insights into the challenges of using of NGS data at other genomic regions of high diversity.\n\nData available in public repositories\n\nhttps://github.com/deboraycb/reliability_hla_1000g

Genomics

Inexpensive Multiplexed Library Preparation for Megabase-Sized Genomes

Whole-genome sequencing has become an indispensible tool of modern biology. However, the cost of sample preparation relative to the cost of sequencing remains high, especially for small genomes where the former is dominant. Here we present a protocol for the rapid and inexpensive preparation of hundreds of multiplexed genomic libraries for Illumina sequencing. By carrying out the Nextera tagmentation reaction in small volumes, replacing costly reagents with cheaper equivalents, and omitting unnecessary steps, we achieve a cost of library preparation of $8 per sample, approximately 6 times cheaper than the widely-used Nextera XT protocol. Furthermore, our procedure takes less than 5 hours for 96 samples and uses nanograms of genomic DNA. Many hundreds of samples can then be pooled on the same HiSeq lane via custom barcodes. Our method is especially useful for re-sequencing of large numbers of full microbial or viral genomes, including those from evolution experiments, genetic screens, and environmental samples.

Genomics

Extensive de novo mutation rate variation between individuals and across the genome of Chlamydomonas reinhardtii

Describing the process of spontaneous mutation is fundamental for understanding the genetic basis of disease, the threat posed by declining population size in conservation biology, and in much evolutionary biology. However, directly studying spontaneous mutation is difficult because of the rarity of de novo mutations. Mutation accumulation (MA) experiments overcome this by allowing mutations to build up over many generations in the near absence of natural selection. In this study, we sequenced the genomes of 85 MA lines derived from six genetically diverse wild strains of the green alga Chlamydomonas reinhardtii. We identified 6,843 spontaneous mutations, more than any other study of spontaneous mutation. We observed seven-fold variation in the mutation rate among strains and that mutator genotypes arose, increasing the mutation rate dramatically in some replicates. We also found evidence for fine-scale heterogeneity in the mutation rate, driven largely by the sequence flanking mutated sites, and by clusters of multiple mutations at closely linked sites. There was little evidence, however, for mutation rate heterogeneity between chromosomes or over large genomic regions of 200Kbp. Using logistic regression, we generated a predictive model of the mutability of sites based on their genomic properties, including local GC content, gene expression level and local sequence context. Our model accurately predicted the average mutation rate and natural levels of genetic diversity of sites across the genome. Notably, trinucleotides vary 17-fold in rate between the most mutable and least mutable sites. Our results uncover a rich heterogeneity in the process of spontaneous mutation both among individuals and across the genome.

Genomics

Distinctive Features of a Saudi Genome

We have fully sequenced the genome of an individual from the region of Saudi Arabia. In order to facilitate comparative analysis, an initial characterization of the new genome was undertaken based on single nucleotide polymorphism (SNP). The SNP data having associated population statistics, essentially the HapMap, served to identify features that were rare by comparison. Methods were developed and applied to tag observed SNPs as different and were extended to identify strings or clusters of difference in the individual relative to comparison populations to effectively increase the significance over single SNP comparison. Difference strings identified in the individual relative to each comparison population showed a genome location pattern with various levels of overlap between the comparison populations. The SNP frequencies from the HapMap population samples Ceu and Yri showed a difference inversion relative to the sample genome. The total SNP difference count was greatest between the individual and the Yri population sample while the number and total span of SNP difference clusters was greatest in comparison with the Ceu population sample. The final pattern of difference clusters has served to define distinctive features in the individual genome toward preliminary characterization.

Genomics

The power of single molecule real-time sequencing technology in the de novo assembly of a eukaryotic genome

Second-generation sequencers (SGS) have been game-changing, achieving cost-effective whole genome sequencing in many non-model organisms. However, a large portion of the genomes still remains unassembled. We reconstructed azuki bean (Vigna angularis) genome using single molecule real-time (SMRT) sequencing technology and achieved the best contiguity and coverage among currently assembled legume crops. The SMRT-based assembly produced 100 times longer contigs with 100 times smaller amount of gaps compared to the SGS-based assemblies. A detailed comparison between the assemblies revealed that the SMRT-based assembly enabled a more comprehensive gene annotation than the SGS-based assemblies where thousands of genes were missing or fragmented. A chromosome-scale assembly was generated based on the high-density genetic map, covering 86% of the azuki bean genome. We demonstrated that SMRT technology, though still needed support of SGS data, achieved a near-complete assembly of a eukaryotic genome.

Genomics

The two-speed genomes of filamentous pathogens: waltz with plants

Fungi and oomycetes include deep and diverse lineages of eukaryotic plant pathogens. The last 10 years have seen the sequencing of the genomes of a multitude of species of these so-called filamentous plant pathogens. Already, fundamental concepts have emerged. Filamentous plant pathogen genomes tend to harbor large repertoires of genes encoding virulence effectors that modulate host plant processes. Effector genes are not randomly distributed across the genomes but tend to be associated with compartments enriched in repetitive sequences and transposable elements. These findings have led to the \"two-speed genome\" model in which filamentous pathogen genomes have a bipartite architecture with gene sparse, repeat rich compartments serving as a cradle for adaptive evolution. Here, we review this concept and discuss how plant pathogens are great model systems to study evolutionary adaptations at multiple time scales. We will also introduce the next phase of research on this topic.

Genomics

The distribution and impact of common copy-number variation in the genome of the domesticated apple, Malus x domestica Borkh.

BackgroundCopy number variation (CNV) is a common feature of eukaryotic genomes, and a growing body of evidence suggests that genes affected by CNV are enriched in processes that are associated with environmental responses. Here we use next generation sequence (NGS) data to detect copy-number variable regions (CNVRs) within the Malus x domestica genome, as well as to examine their distribution and impact.\n\nMethodsCNVRs were detected using NGS data derived from 30 accessions of M. x domestica analyzed using the read-depth method, as implemented in the CNVrd2 software. To improve the reliability of our results, we developed a quality control and analysis procedure that involved checking for organelle DNA, not repeat masking, and the determination of CNVR identity using a permutation testing procedure.\n\nResultsOverall, we identified 876 CNVRs, which spanned 3.5% of the apple genome. To verify that detected CNVRs were not artifacts, we analyzed the B-allele-frequencies (BAF) within a single nucleotide polymorphism (SNP) array dataset derived from a screening of 185 individual apple accessions and found the CNVRs were enriched for SNPs having aberrant BAFs (P < 1e-13, Fishers Exact test). Putative CNVRs overlapped 845 gene models and were enriched for resistance (R) gene models (P < 1e-22, Fishers exact test). Of note was a cluster of resistance gene models on chromosome 2 near a region containing multiple major gene loci conferring resistance to apple scab.\n\nConclusionWe present the first analysis and catalogue of CNVRs in the M. x domestica genome. The enrichment of the CNVRs with R gene models and their overlap with gene loci of agricultural significance draw attention to a form of unexplored genetic variation in apple. This research will underpin further investigation of the role that CNV plays within the apple genome.

Genomics

Whole genome sequence analyses of Western Central African Pygmy hunter-gatherers reveal a complex demographic history and identify candidate genes under positive natural selection

African Pygmies practicing a mobile hunter-gatherer lifestyle are phenotypically and genetically diverged from other anatomically modern humans, and they likely experienced strong selective pressures due to their unique lifestyle in the Central African rainforest. To identify genomic targets of adaptation, we sequenced the genomes of four Biaka Pygmies from the Central African Republic and jointly analyzed these data with the genome sequences of three Baka Pygmies from Cameroon and nine Yoruba famers. To account for the complex demographic history of these populations that includes both isolation and gene flow, we fit models using the joint allele frequency spectrum and validated them using independent approaches. Our two best-fit models both suggest ancient divergence between the ancestors of the farmers and Pygmies, 90,000 or 150,000 years ago. We also find that bi-directional asymmetric gene-flow is statistically better supported than a single pulse of unidirectional gene flow from farmers to Pygmies, as previously suggested. We then applied complementary statistics to scan the genome for evidence of selective sweeps and polygenic selection. We found that conventional statistical outlier approaches were biased toward identifying candidates in regions of high mutation or low recombination rate. To avoid this bias, we assigned P-values for candidates using whole-genome simulations incorporating demography and variation in both recombination and mutation rates. We found that genes and gene sets involved in muscle development, bone synthesis, immunity, reproduction, cell signaling and development, and energy metabolism are likely to be targets of positive natural selection in Western African Pygmies or their recent ancestors.

Genomics

The Nicrophorus vespilloides genome and methylome, a beetle with complex social behavior

Testing for conserved and novel mechanisms underlying phenotypic evolution requires a diversity of genomes available for comparison spanning multiple independent lineages. For example, complex social behavior in insects has been investigated primarily with eusocial lineages, nearly all of which are Hymenoptera. If conserved genomic influences on sociality do exist, we need data from a wider range of taxa that also vary in their levels of sociality. Here we present information on the genome of the subsocial beetle Nicrophorus vespilloides, a species long used to investigate evolutionary questions of complex social behavior. We used this genome to address two questions. First, does life history predict overlap in gene models more strongly than phylogenetic groupings? Second, like other insects with highly developed social behavior but unlike other beetles, does N. vespilloides have DNA methylation? We found the overlap in gene models was similar between N. vespilloides and all other insect groups regardless of life history. Unlike previous studies of beetles, we found strong evidence of DNA methylation, which allows this species to be used to address questions about the potential role of methylation in social behavior. The addition of this genome adds a coleopteran resource to answer questions about the evolution and mechanistic basis of sociality.

Genomics

Conversion of Genomic DNA to Proxy Constructs Suitable for Accurate Nanopore Sequencing

Nanopore sequencing at single-base resolution is challenging. There are developing technologies to convert DNA molecules to expanded constructs. Such constructs can be sequenced by nanopores in place of the original DNA molecules. We present a novel method for converting genomic DNA to expanded constructs (\"proxies\") with 99.67% accuracy. Our method \"reads\" each base in each DNA fragment and appends an oligonucleotide to the DNA fragment after each base \"reading\". Each appended oligonucleotide represents a specific base type, so that the proxy construct consisting of all the appended oligonucleotides faithfully represents the original DNA sequence. We generated proxies for genomic DNA and confirmed the identities of both the proxies and their corresponding original DNA sequences by performing sequencing using Ion Torrent sequencer.\n\nConversion to proxies had only 0.33% raw error rate. Errors were: 93.96% deletions, 5.29% insertions, and 0.74% substitutions. The longest sequenced proxy was 170 bases, corresponding to a 17-base original DNA sequence. The short length of the detected proxies reflected restrictions imposed by Ion Torrents short reads and was not caused by limitations of our method. The consensus sequence built by using proxies alone (average length: 120 bases; corresponding to original sequences with average length 12 bases) covered 55% of the reference genome with 100% accuracy, and outperformed the Ion Torrent sequencing of the corresponding original DNA fragments in terms of accuracy, coverage and number of aligned sequences. Data and other materials can be found at http://www.vastogen.com/data.html. This proof-of-concept experiment demonstrates highly accurate proxy construction at the whole genome level. To our knowledge, this is the first demonstrated construction of expanded versions of DNA at the whole genome level.

Genomics

FecalSeq: methylation-based enrichment for noninvasive population genomics from feces

Obtaining high-quality samples from wild animals is a major obstacle for genomic studies of many taxa, particular at the population level, as collection methods for such samples are typically invasive. DNA from feces is easy to obtain noninvasively, but is dominated by a preponderance of bacterial and other non-host DNA. Because next-generation sequencing technology sequences DNA largely indiscriminately, the high proportion of exogenous DNA drastically reduces the efficiency of high-throughput sequencing for host animal genomics. In order to address this issue, we developed an inexpensive methylation-based capture method for enriching host DNA from noninvasively obtained fecal DNA samples. Our method exploits natural differences in CpG-methylation density between vertebrate and bacterial genomes to preferentially bind and isolate host DNA from majority-bacterial fecal DNA samples. We demonstrate that the enrichment is robust, efficient, and compatible with downstream library preparation methods useful for population studies (e.g., RADseq). Compared to other enrichment strategies, our method is quick and inexpensive, adding only a negligible cost to sample preparation for research that is often severely constrained by budgetary limitations. In combination with downstream methods such as RADseq, our approach allows for cost-effective and customizable genomic-scale genotyping that was previously feasible in practice only with invasive samples. Because feces are widely available and convenient to collect, our method empowers researchers to explore genomic-scale population-level questions in organisms for which invasive sampling is challenging or undesirable.

Genomics

No evidence for extensive horizontal gene transfer in the genome of the tardigrade Hypsibius dujardini

Tardigrades are meiofaunal ecdysozoans that are key to understanding the origins of Arthropoda. Many species of Tardigrada can survive extreme conditions through cryptobiosis. In a recent paper (Boothby TC et al (2015) Evidence for extensive horizontal gene transfer from the draft genome of a tardigrade. Proc Natl Acad Sci USA 112:15976-15981) the authors concluded that the tardigrade Hypsibius dujardini had an unprecedented proportion (17%) of genes originating through functional horizontal gene transfer (fHGT), and speculated that fHGT was likely formative in the evolution of cryptobiosis. We independently sequenced the genome of H. dujardini. As expected from whole-organism DNA sampling, our raw data contained reads from non-target genomes. Filtering using metagenomics approaches generated a draft H. dujardini genome assembly of 135 Mb with superior assembly metrics to the previously published assembly. Additional microbial contamination likely remains. We found no support for extensive fHGT. Among 23,021 gene predictions we identified 0.2% strong candidates for fHGT from bacteria, and 0.2% strong candidates for fHGT from non-metazoan eukaryotes. Cross-comparison of assemblies showed that the overwhelming majority of HGT candidates in the Boothby et al. genome derived from contaminants. We conclude that fHGT into H. dujardini accounts for at most 1-2% of genes and that the proposal that one sixth of tardigrade genes originate from functional HGT events is an artefact of undetected contamination.

Genomics

Clinical Validation of a Non-Invasive Prenatal Test for Genome-Wide Detection of Fetal Copy Number Variants

1.BackgroundCurrent cell-free DNA (cfDNA) assessment of fetal chromosomes does not analyze and report on all chromosomes. Hence, a significant proportion of fetal chromosomal abnormalities are not detectable by current non-invasive methods. Here we report the clinical validation of a novel NIPT designed to detect genome-wide gains and losses of chromosomal material [&ge;]7 Mb and losses associated with specific deletions <7 Mb.\n\nObjectiveThe objective of this study is to provide a clinical validation of the sensitivity and specificity of a novel NIPT for detection of genome-wide abnormalities.\n\nStudy DesignThis retrospective, blinded study included maternal plasma collected from 1222 study subjects with pregnancies at increased risk for fetal chromosomal abnormalities that were assessed for trisomy 21 (T21), trisomy 18 (T18), trisomy 13 (T13), sex chromosome aneuploidies (SCAs), fetal sex, genome-wide copy number variants (CNVs) 7 Mb and larger, and select deletions smaller than 7 Mb. Performance was assessed by comparing test results with findings from G-band karyotyping, microarray data, or high coverage sequencing.\n\nResultsClinical sensitivity within this study was determined to be 100% for T21, T18, T13, and SCAs, and 97.7% for genome-wide CNVs. Clinical specificity within this study was determined to be 100% for T21, T18, and T13, and 99.9% for SCAs and CNVs. Fetal sex classification had an accuracy of 99.6%.\n\nConclusionThis study has demonstrated that genome-wide non-invasive prenatal testing (NIPT) for fetal chromosomal abnormalities can provide high resolution, sensitive, and specific detection of a wide range of sub-chromosomal and whole chromosomal abnormalities that were previously only detectable by invasive karyotype analysis. In some instances, this NIPT also provided additional clarification about the origin of genetic material that had not been identified by invasive karyotype analysis.

Genomics