bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Genome-wide Associations of Flavivirus Capsid Proteins

Dengue virus (DENV) and Zika virus (ZIKV) are both positive sense single-stranded RNA viruses. They are packaged within the virion with a capsid (C) protein to form the nucleocapsid. Based on cryo-electron microscopy imaging, the nucleocapsid has been described as lacking symmetry, whilst there is distinguishable separation of the C proteins from the viral RNA (vRNA) genome. Here, to elucidate the architecture of the nucleocapsid of DENV serotype 2 and ZIKV, we used a nuclease digestion assay and next-generation sequencing to map the respective vRNA genome wide association with the C protein in vitro. The C protein exhibited non-uniform binding along the vRNA, and as C protein concentration increased, the normalized read counts also increased. A saturation point of 1:100 (vRNA:C protein monomers) was found, and binding regions showed variable saturation patterns. We also observed that C protein had a preference for G-rich sequences for both viruses. Taken together, we demonstrate that the DENV 2 and ZIKV C proteins bind vRNA in a non-uniform manner with distinct patterns of association.\n\nSingificance StatementOur study demonstrates that flavivirus capsid proteins associate with the viral genome at specific sites rather than in a uniform manner as commonly expected. We estimate the number of capsid proteins binding to a single genomic RNA. We proceed to locate the capsid binding sites along the viral genomes of Dengue and Zika viruses. We characterize the binding sites in terms of affinity and analyze the nucleotide composition and sequence motifs at binding sites. We cross-reference binding sites against SHAPE reactivity data corresponding to local RNA secondary structure, which allows us to identify structural motifs of capsid binding sites. As capsid proteins are essential for viral packaging, these interactions may form attractive targets for therapeutic intervention.

genomics

CNCC: An analysis tool to determine genome-wide DNA break end structure at single-nucleotide resolution

DNA double-stranded breaks (DSBs) are potentially deleterious events in a cell. The end structures (blunt, 3- and 5-overhangs) at sites of double-stranded breaks contribute to the fate of their repair and provide critical information for consequences of the damage. Here, we describe the use of a coverage-normalized cross correlation analysis (CNCC) to process high-precision genome-wide break mapping data, and determine genome-wide break end structure distributions at single-nucleotide resolution. For the first time, on a genome-wide scale, our analysis revealed the increase in the 5 to 3 end resection following etoposide treatment, and the global progression of the resection due to the removal of DNA topoisomerase II cleavage complexes. Further, our method distinguished the change in the pattern of DSB end structure with increasing doses of the drug. The ability of this method to determine DNA break end structures without a priori knowledge of break sequences or genomic position should have broad applications in understanding genome instability.

genomics

The Chinese chestnut genome: a reference for species restoration

Forest tree species are increasingly subject to severe mortalities from exotic pests, diseases, and invasive organisms, accelerated by climate change. Forest health issues are threatening multiple species and ecosystem sustainability globally. While sources of resistance may be available in related species, or among surviving trees, introgression of resistance genes into threatened tree species in reasonable time frames requires genome-wide breeding tools. Asian species of chestnut (Castanea spp.) are being employed as donors of disease resistance genes to restore native chestnut species in North America and Europe. To aid in the restoration of threatened chestnut species, we present the assembly of a reference genome with chromosome-scale sequences for Chinese chestnut (C. mollissima), the disease-resistance donor for American chestnut restoration. We also demonstrate the value of the genome as a platform for research and species restoration, including new insights into the evolution of blight resistance in Asian chestnut species, the locations in the genome of ecologically important signatures of selection differentiating American chestnut from Chinese chestnut, the identification of candidate genes for disease resistance, and preliminary comparisons of genome organization with related species.

genomics

Identification of Pathogenic Structural Variants in Rare Disease Patients through Genome Sequencing

PurposeClinical whole genome sequencing is becoming more common for determining the molecular diagnosis of rare disease. However, standard clinical practice often focuses on small variants such as single nucleotide variants and small insertions/deletions. This leaves a wide range of larger \"structural variants\" that are not commonly analyzed in patients.\n\nMethodsWe developed a pipeline for processing structural variants for patients who received whole genome sequencing through the Undiagnosed Diseases Network (UDN). This pipeline called structural variants, stored them in an internal database, and filtered the variants based on internal frequencies and external annotations. The remaining variants were manually inspected and then interesting findings were reported as research variants to clinical sites in the UDN.\n\nResultsOf 477 analyzed UDN cases, 286 cases ({approx} 60%) received at least one structural variant as a research finding. The variants in 16 cases ({approx} 4%) are considered \"Certain\" or \"Highly likely\" molecularly diagnosed and another 4 cases are currently in review. Of those 20 cases, at least 13 were identified originally through our pipeline with one finding leading to identification of a new disease. As part of this paper, we have also released the collection of variant calls identified in our cohort along with heterozygous and homozygous call counts. This data is available at https://github.com/HudsonAlpha/UDN_SV_export.\n\nConclusionStructural variants are key genetic features that should be analyzed during routine clinical genomic analysis. For our UDN patients, structural variants helped solve {approx} 4% of the total number of cases ({approx} 13% of all genome sequencing solves), a success rate we expect to improve with better tools and greater understanding of the human genome.

genomics

Genomic sequencing of the aquatic Fusarium spp. QHM and BWC1 and their potential application in environmental protection

Fusarium species are distributed widely in ecosystems of a wide pH range and play a pivotal role in the aquatic community through the degradation of xenobiotic compounds and secretion of secondary metabolites. The elucidation of their genome would therefore be highly impactful with regard to the control of environmental pollution. Therefore, in this study, two indigenous strains of aquatic Fusarium, QHM and BWC1, were isolated from a coal mine pit and a subterranean river respectively, cultured under acidic conditions, and sequenced. Phylogenetic analysis of these two isolates was conducted based on the sequences of internal transcript (ITS1 and ITS4) and encoding {beta}-microtubulin (TUB2), translation elongation factors (TEFs) and the second large sub-unit of RNA polymerase (RPB2). Fusarium, QHM could potentially represent a new species within the Fusarium fujikuroi species complex. Fusarium BWC1 were found to form a clade with Fusarium subglutinans NRRL 22016, and predicted to be Fusarium subglutinans. Shot-gun sequencing on the Illumina Hiseqx10 Platform was used to elucidate the draft genomes of the two species. Gene annotation and functional analyses revealed that they had bio-degradation pathways for aromatic compounds; further, their main pathogenic mechanism was found to be the efflux pump. To date, the genomes of only a limited number of acidic species from the Fusarium fujikuroi species complex, especially from the aquatic species, have been sequenced. Therefore, the present findings are novel and have important potential for the future in terms of environmental control.\n\nIMPORTANCEFusarium genus has over 300 species and were distributed in a variety of ecosystem. Increasing attention has been drawn to Fusarium due to the importance in aquatic community, pathogenicity and environmental protection. The genomes of the strains in this work isolated in acidic condition, were sequenced. The analysis has indicated that the isolates were able to biodegrade xenobiotics, which makes it potentially function as environmental bio-agent for aromatic pollution control and remediation. Meanwhile, the virulence and pathogenicity were also predicted for reference of infection control. The genome information may lay foundation for the fungal identification, disease prevention resulting from these isolates and other \"-omics\" research. The isolates were phylogenetically classified into Fusarium fujikuroi species complex by means of concatenated gene analysis, serving as new addition to the big complex.

genomics

Aquila: diploid personal genome assembly and comprehensive variant detection based on linked reads

Variant discovery in personal, whole genome sequence data is critical for uncovering the genetic contributions to health and disease. We introduce a new approach, Aquila, that uses linked-read data for generating a high quality diploid genome assembly, from which it then comprehensively detects and phases personal genetic variation. Assemblies cover >95% of the human reference genome, with over 98% in a diploid state. Thus, the assemblies support detection and accurate genotyping of the most prevalent types of human genetic variation, including single nucleotide polymorphisms (SNPs), small insertions and deletions (small indels), and structural variants (SVs), in all but the most difficult regions. All heterozygous variants are phased in blocks that can approach arm-level length. The final output of Aquila is a diploid and phased personal genome sequence, and a phased VCF file that also contains homozygous and a few unphased heterozygous variants. Aquila represents a cost-effective evolution of whole-genome reconstruction that can be applied to cohorts for variation discovery or association studies, or to single individuals with rare phenotypes that could be caused by SVs or compound heterozygosity.

genomics

Chromosome level draft genomes of the fall armyworm, Spodoptera frugiperda (Lepidoptera: Noctuidae), an alien invasive pest in China

The fall armyworm (FAW), Spodoptera frugiperda (J.E. Smith) is a severely destructive pest native to the Americas, but has now become an alien invasive pest in China, and causes significant economic loss. Therefore, in order to make effective management strategies, it is highly essential to understand genomic architecture and its genetic background. In this study, we assembled two chromosome scale genomes of the fall armyworm, representing one male and one female individual procured from Yunnan province of China. The genome sizes were identified as 542.42 Mb with N50 of 14.16 Mb, and 530.77 Mb with N50 of 14.89 Mb for the male and female FAW, respectively. We predicted about 22,201 genes in the male genome. We found the expansion of cytochrome P450 and glutathione s-transferase gene families, which are functionally related to the intensified detoxification and pesticides tolerance. Further population analyses of corn strain (C strain) and rice strain (R strain) revealed that the Chinese fall armyworm was most likely invaded from Africa. These strain information, genome features and possible invasion source described in this study will be extremely important for making effective strategies to manage the fall armyworms.

genomics

Chromosome-scale assembly comparison of the Korean Reference Genome KOREF from PromethION and PacBio with Hi-C mapping information

BackgroundLong DNA reads produced by single molecule and pore-based sequencers are more suitable for assembly and structural variation discovery than short read DNA fragments. For de novo assembly, PacBio and Oxford Nanopore Technologies (ONT) are favorite options. However, PacBios SMRT sequencing is expensive for a full human genome assembly and costs over 40,000 USD for 30x coverage as of 2019. ONT PromethION sequencing, on the other hand, is one-twelfth the price of PacBio for the same coverage. This study aimed to compare the cost-effectiveness of ONT PromethION and PacBios SMRT sequencing in relation to the quality.\n\nFindingsWe performed whole genome de novo assemblies and comparison to construct an improved version of KOREF, the Korean reference genome, using sequencing data produced by PromethION and PacBio. With PromethION, an assembly using sequenced reads with 64x coverage (193 Gb, 3 flowcell sequencing) resulted in 3,725 contigs with N50s of 16.7 Mbp and a total genome length of 2.8 Gbp. It was comparable to a KOREF assembly constructed using PacBio at 62x coverage (188 Gbp, 2,695 contigs and N50s of 17.9 Mbp). When we applied Hi-C-derived long-range mapping data, an even higher quality assembly for the 64x coverage was achieved, resulting in 3,179 scaffolds with an N50 of 56.4 Mbp.\n\nConclusionThe pore-based PromethION approach provides a good quality chromosome-scale human genome assembly at a low cost with long maximum contig and scaffold lengths and is more cost-effective than PacBio at comparable quality measurements.

genomics

Insights into human genetic variation and population history from 929 diverse genomes

Genome sequences from diverse human groups are needed to understand the structure of genetic variation in our species and the history of, and relationships between, different populations. We present 929 high-coverage genome sequences from 54 diverse human populations, 26 of which are physically phased using linked-read sequencing. Analyses of these genomes reveal an excess of previously undocumented private genetic variation in southern and central Africa and in Oceania and the Americas, but an absence of fixed, private variants between major geographical regions. We also find deep and gradual population separations within Africa, contrasting population size histories between hunter-gatherer and agriculturalist groups in the last 10,000 years, a potentially major population growth episode after the peopling of the Americas, and a contrast between single Neanderthal but multiple Denisovan source populations contributing to present-day human populations. We also demonstrate benefits to the study of population relationships of genome sequences over ascertained array genotypes. These genome sequences are freely available as a resource with no access or analysis restrictions.

genomics

A classification framework for Bacillus anthracis defined by global genomic structure

Bacillus anthracis, the causative agent of anthrax, is a considerable global health threat affecting wildlife, livestock, and the general public. In this study whole-genome sequence analysis of over 350 B. anthracis isolates was used to establish a new high-resolution global genotyping framework that is both biogeographically informative, and compatible with multiple genomic assays. The data presented in this study shed new light on the diverse global dissemination of this species and indicate that many lineages may be uniquely suited to the geographic regions in which they are found. In addition, we demonstrate that plasmid genomic structure for this species is largely consistent with chromosomal population structure, suggesting vertical inheritance in this bacterium has contributed to its evolutionary persistence. This classification methodology is the first based on population genomic structure for this species and has potential use for local and broader institutions seeking to understand both disease outbreak origins and recent introductions. In addition, we provide access to a newly developed genotyping script as well as the full whole genome sequence analyses output for this study, allowing future studies to rapidly employ and append their data in the context of this global collection. This framework may act as a powerful tool for public health agencies, wildlife disease laboratories, and researchers seeking to utilize and expand this classification scheme for further investigations into B. anthracis evolution.

genomics

Concurrent Genome and Epigenome Editing by CRISPR-Mediated Sequence Replacement

Recent advances in genome editing have facilitated the direct manipulation of not only the genome, but also the epigenome. Genome editing is typically performed by introducing a single CRISPR/Cas9-mediated double stranded break (DSB), followed by NHEJ or HDR mediated repair. Epigenome editing, and in particular methylation of CpG dinucleotides, can be performed using catalytically inactive Cas9 (dCas) fused to a methyltransferase domain. However, for investigations of the role of methylation in gene silencing, studies based on dCas9-methyltransferase have limited resolution and are potentially confounded by the effects of binding of the fusion protein. As an alternative strategy for epigenome editing, we tested CRISPR/Cas9 dual cutting of the genome in the presence of in vitro methylated exogenous DNA, i.e. to drive replacement of the DNA sequence intervening the dual cuts via NHEJ. In a proof-of-concept at the HPRT1 promoter, successful replacement events with heavily methylated alleles of a CpG island resulted in functional silencing of the HPRT1 gene. Although still limited in efficiency, our study demonstrates concurrent epigenome and genome editing in a single event, and opens the door to investigations of the functional consequences of methylation patterns at single CpG dinucleotide resolution. Our results furthermore support the conclusion that promoter methylation is sufficient to functionally silence gene expression.

genomics

Accurate assembly of the olive baboon (Papio anubis) genome using long-read and Hi-C data

Besides macaques, baboons are the most commonly used nonhuman primate in biomedical research. Despite this importance, the genomic resources for baboons are quite limited. In particular, the current baboon reference genome Panu_3.0 is a highly fragmented, reference-guided (i.e., not fully de novo) assembly, and its poor quality inhibits our ability to conduct downstream genomic analyses. Here we present a truly de novo genome assembly of the olive baboon (Papio anubis) that uses data from several recently developed single-molecule technologies. Our assembly, Panubis1.0, has an N50 contig size of ~1.46 Mb (as opposed to 139 Kb for Panu_3.0), and has single scaffolds that span each of the 20 autosomes and the X chromosome. We highlight multiple lines of evidence (including Bionano Genomics data, pedigree linkage information, and linkage disequilibrium data) suggesting that there are several large assembly errors in Panu_3.0, which have been corrected in Panubis1.0.

genomics

RADICL-seq identifies general and cell type-specific principles of genome-wide RNA-chromatin interactions

Mammalian genomes encode tens of thousands of noncoding RNAs. Most noncoding transcripts exhibit nuclear localization and several have been shown to play a role in the regulation of gene expression and chromatin remodelling. To investigate the function of such RNAs, methods to massively map the genomic interacting sites of multiple transcripts have been developed. However, they still present some limitations. Here, we introduce RNA And DNA Interacting Complexes Ligated and sequenced (RADICL-seq), a technology that maps genome-wide RNA-chromatin interactions in intact nuclei. RADICL-seq is a proximity ligation-based methodology that reduces the bias for nascent transcription, while increasing genomic coverage and unique mapping rate efficiency compared to existing methods. RADICL-seq identifies distinct patterns of genome occupancy for different classes of transcripts as well as cell type-specific RNA-chromatin interactions, and emphasizes the role of transcription in the establishment of chromatin structure.

genomics

Strains used in whole organism Plasmodium falciparum vaccine trials differ in genome structure, sequence, and immunogenic potential

BackgroundPlasmodium falciparum (Pf) whole-organism sporozoite vaccines have provided excellent protection against controlled human malaria infection (CHMI) and naturally transmitted heterogeneous Pf in the field. Initial CHMI studies showed significantly higher durable protection against homologous than heterologous strains, suggesting the presence of strain-specific vaccine-induced protection. However, interpretation of these results and understanding of their relevance to vaccine efficacy (VE) have been hampered by the lack of knowledge on genetic differences between vaccine and CHMI strains, and how these strains are related to parasites in malaria endemic regions.\n\nMethodsWhole genome sequencing using long-read (Pacific Biosciences) and short-read (Illumina) sequencing platforms was conducted to generate de novo genome assemblies for the vaccine strain, NF54, and for strains used in heterologous CHMI (7G8 from Brazil, NF166.C8 from Guinea, and NF135.C10 from Cambodia). The assemblies were used to characterize sequence polymorphisms and structural variants in each strain relative to the reference Pf 3D7 (a clone of NF54) genome. Strains were compared to each other and to clinical isolates from South America, Sub-Saharan Africa, and Southeast Asia.\n\nResultsWhile few variants were detected between 3D7 and NF54, we identified tens of thousands of variants between NF54 and the three heterologous strains both genome-wide and within regulatory and immunologically important regions, including in pre-erythrocytic antigens that may be key for sporozoite vaccine-induced protection. Additionally, these variants directly contribute to diversity in immunologically important regions of the genomes as detected through in silico CD8+ T cell epitope predictions. Of all heterologous strains, NF135.C10 consistently had the highest number of unique predicted epitope sequences when compared to NF54, while NF166.C8 had the lowest. Comparison to global clinical isolates revealed that these four strains are representative of their geographic region of origin despite long-term culture adaptation; of note, NF135.C10 is from an admixed population, and not part of recently-formed drug resistant subpopulations present in the Greater Mekong Sub-region.\n\nConclusionsThese results are assisting the interpretation of VE of whole-organism vaccines against homologous and heterologous CHMI, and may be useful in informing the choice of strains for inclusion in region-specific or multi-strain vaccines.

genomics

Long-read assembly of the Chinese rhesus macaque genome and identification of ape-specific structural variants

Rhesus macaque (Macaca mulatta) is a widely-studied nonhuman primate. Here we present a high-quality de novo genome assembly of the Chinese rhesus macaque (rheMacS) using long-read sequencing and multiplatform scaffolding approaches. Compared to the current Indian rhesus macaque reference genome (rheMac8), the rheMacS genome assembly improves sequence contiguity by 75-fold, closing 21,940 of the remaining assembly gaps (60.8 Mbp). To improve gene annotation, we generated more than two million full-length transcripts from ten different tissues by long-read RNA sequencing. We sequence resolve 53,916 structural variants (96% novel) and identify 17,000 ape-specific structural variants (ASSVs) based on comparison to the long-read assembly of ape genomes. We show that many ASSVs map within ChIP-seq predicted enhancer regions where apes and macaque show diverged enhancer activity and gene expression. We further characterize a set of candidate ASSVs that may contribute to ape- or great-ape-specific phenotypic traits, including taillessness, brain volume expansion, improved manual dexterity, and large body size. This improved rheMacS genome assembly serves as an ideal reference for future biomedical and evolutionary studies.

genomics

Comparative genomics of Alternaria species provides insights into the pathogenic lifestyle of Alternaria brassicae - a pathogen of the Brassicaceae family

Alternaria brassicae, a necrotrophic pathogen, causes Alternaria Leaf Spot, one of the economically important diseases of Brassica crops. Many other Alternaria spp. such as A. brassicicola and A. alternata are known to cause secondary infections in the A. brassicae-infected Brassicas. The genome architecture, pathogenicity factors, and determinants of host-specificity of A. brassicae are unknown. In this study, we annotated and characterised the recently announced genome assembly of A. brassicae and compared it with other Alternaria spp. to gain insights into its pathogenic lifestyle. Additionally, we sequenced the genomes of two A. alternata isolates that were co-infecting B. juncea. Genome alignments within the Alternaria spp. revealed high levels of synteny between most chromosomes with some intrachromosomal rearrangements. We show for the first time that the genome of A. brassicae, a large-spored Alternaria species, contains a dispensable chromosome. We identified 460 A. brassicae-specific genes, which included many secreted proteins and effectors. Furthermore, we have identified the gene clusters responsible for the production of Destruxin-B, a known pathogenicity factor of A. brassicae. The study provides a perspective into the unique and shared repertoire of genes within the Alternaria genus and identifies genes that could be contributing to the pathogenic lifestyle of A. brassicae.

genomics

Revisiting the landscape of evolutionary breakpoints across human genome using multi-way comparison

Genome rearrangement is one of the major forces driving the processes of the evolution and disease development. The chromosomal position affected by these rearrangements are called breakpoints. The breakpoints occurring during the evolution of species are known to be non randomly distributed. Detecting their landscape and mapping them to genomic features constitute an important features in both comparative and functional genomics. Several studies have attempted to provide such mapping based on pairwise comparison of genes as conservation anchors. With the availability of more accurate multi-way alignments, we design an approach to identify synteny blocks and evolutionary breakpoints based on UCSC 45-way conservation sequence alignments with 12 selected species. The multi-way designed approach with the mild flexibility of presence of selected species, helped to have a better determination of human lineage-specific evolutionary breakpoints. We identified 261,391 human lineage-specific evolutionary breakpoints across the genome and 2,564 dense regions enriched with biological processes involved in adaptive traits such as response to DNA damage stimulus, cellular response to stress and metabolic process. Moreover, we found 230 regions refractory to evolutionary breakpoints that carry genes associated with crucial developmental process such as organ morphogenesis, skeletal system development, chordate embryonic development, nerve development and regulation of biological process. This initial map of the human genome will help to gain better insight into several studies including developmental studies and cancer rearrangement processes.

genomics

Sequencing, de novo assembly and annotation of the genome of the scleractinian coral, Pocillopora acuta

Coral reefs are the most divers marine ecosystem. However, under the pressure of global changes and anthropogenic disturbances corals and coral reefs are declining worldwide. In order to better predict and understand the future of these organisms all the tools of modern biology are needed today. However, many NGS based approaches are not feasible in corals because of the lack of reference genomes. Therefore we have sequenced, de novo assembled, and annotated, the draft genome of one of the most studied coral species, Pocillopora acuta (ex damicornis). The sequencing strategy was based on four libraries with complementary insert size and sequencing depth (180pb, 100x; 3Kb, 25x; 8kb, 12x and 20 kb, 12x). The de novo assembly was performed with Platanus (352 Mb; 25,553 scaffolds; N50 171,375 bp). 36,140 genes were annotated by RNA-seq data and 64,558 by AUGUSTUS (Hidden-Markov model). Gene functions were predicted through Blast and orthology based approaches. This new genomic resource will enable the development of a large array of genome wide studies but also shows that the de novo assembly of a coral genome is now technically feasible and economically realistic.

genomics