bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

From Estimation to Prediction of Genomic Variances: Allowing for Linkage Disequilibrium and Unbiasedness

The additive genomic variance in linear models with random marker effects can be defined as a random variable that is in accordance with classical quantitative genetics theory. Common approaches to estimate the genomic variance in random-effects linear models based on genomic marker data can be regarded as the unconditional (or prior) expectation of this random additive genomic variance, and result in a negligence of the contribution of linkage disequilibrium. We introduce a novel best prediction (BP) approach for the additive genomic variance in both the current and the base population in the framework of genomic prediction using the gBLUP-method. The resulting best predictor is the conditional (or posterior) expectation of the additive genomic variance when using the additional information given by the phenotypic data, and is structurally in accordance with the genomic equivalent of the classical additive genetic variance in random-effects models. In particular, the best predictor includes the contribution of (marker) linkage disequilibrium to the additive genomic variance and eliminates the missing contribution of LD that is caused by the assumptions of statistical frameworks such as the random-effects model. We derive an empirical best predictor (eBP) and compare its performance with common approaches to estimate the additive genomic variance in random-effects models on commonly used genomic datasets.

genomics

Expanded view of the ecological genomics of ant responses to climate change

Given the abundance, broad distribution, and diversity of roles that ants play in many ecosystems, they are an ideal group to serve as ecosystem indicators of climatic change. At present, only a few whole-genome sequences of ants are available (19 of > 16,000 species), mostly from tropical and sub-tropical species. To address this limited sampling, we sequenced genomes of temperate-latitude species from the genus Aphaenogaster, a genus with important seed dispersers. In total, we sampled seven colonies of six species: A. ashmeadi, A. floridana, A. fulva, A. miamiana, A. picea, and A. rudis. The geographic ranges of these species collectively span eastern North America from southern Florida to southern Canada, which encompasses a latitudinal gradient in which many climatic variables are changing rapidly. For the six genomes, we assembled an average of 271,039 contigs into 47,337 scaffolds. The mean genome size was 370.5 Mb, ranging from 310.3 to 429.7, which is comparable to that of other sequenced ant genomes (212.8 to 396.0 Mb) and flow cytometry estimates (210.7 to 690.4 Mb). In an analysis of currently sequenced ant genomes and the new Aphaenogaster sequences, we found that after controlling for both spatial autocorrelation and phylogenetics ant genome size was marginally correlated with sample site climate similarity. Of all examined climate variables, minimum temperature showed the strongest correlation with genome size, with ants from locations with colder minimum temperatures having larger genomes. These results suggest that temperature extremes could be a selective force acting on ant genomes and point to the need for more extensive sequencing of ant genomes.

genomics

Haplotype-resolved and integrated genome analysis of ENCODE cell line HepG2

The HepG2 cancer cell line is one of the most widely-used biomedical research and one of the main cell lines of ENCODE. Vast numbers of functional genomics and epigenomics datasets have been produced to characterize its biology. However, the correct interpretation such data requires an understanding of the cell lines genome sequence and genome structure. Using a variety of sequencing and analysis methods, we identified a wide spectrum of HepG2 genome characteristics: copy numbers of chromosomal segments, SNVs and Indels (corrected for aneuploidy), phased haplotypes extending to entire chromosome arms, loss of heterozygosity, retrotransposon insertions, structural variants (SVs) including complex and somatic genomic rearrangements. We also identified allele-specific expression and DNA methylation genome-wide and assembled an allele-specific CRISPR/Cas9 targeting map.\n\nSIGNIFICANCEHaplotype-resolved and comprehensive whole-genome analysis of a widely-used cell line for cancer research and ENCODE, HepG2, serves as an essential resource for unlocking complex cancer gene regulation using a genome-integrated framework and also provides genomic context for the analysis of ~1,000 functional datasets to date on ENCODE for biological discovery. We also demonstrate how deeper insights into genomic regulatory complexity are gained by adopting a genome-integrated framework.

genomics

ML-DSP: Machine Learning with Digital Signal Processing for ultrafast, accurate, and scalable genome classification at all taxonomic levels

BackgroundAlthough methods and software tools abound for the comparison, analysis, identification, and taxonomic classification of the enormous amount of genomic sequences that are continuously being produced, taxonomic classification remains challenging. The difficulty lies within both the magnitude of the dataset and the intrinsic problems associated with classification. The need exists for an approach and software tool that addresses the limitations of existing alignment-based methods, as well as the challenges of recently proposed alignment-free methods.\n\nResultsWe combine supervised Machine Learning with Digital Signal Processing to design ML-DSP, an alignment-free software tool for ultrafast, accurate, and scalable genome classification at all taxonomic levels.\n\nWe test ML-DSP by classifying 7,396 full mitochondrial genomes from the kingdom to genus levels, with 98% classification accuracy. Compared with the alignment-based classification tool MEGA7 (with sequences aligned with either MUSCLE, or CLUSTALW), ML-DSP has similar accuracy scores while being significantly faster on two small benchmark datasets (2,250 to 67,600 times faster for 41 mammalian mitochondrial genomes). ML-DSP also successfully scales to accurately classify a large dataset of 4,322 complete vertebrate mtDNA genomes, a task which MEGA7 with MUSCLE or CLUSTALW did not complete after several hours, and had to be terminated. ML-DSP also outperforms the alignment-free tool FFP (Feature Frequency Profiles) in terms of both accuracy and time, being three times faster for the vertebrate mtDNA genomes dataset.\n\nConclusionsWe provide empirical evidence that ML-DSP distinguishes complete genome sequences at all taxonomic levels. Ultrafast and accurate taxonomic classification of genomic sequences is predicted to be highly relevant in the classification of newly discovered organisms, in distinguishing genomic signatures, in identifying mechanistic determinants of genomic signatures, and in evaluating genome integrity.

genomics

Fitness Landscape of the Fission Yeast Genome

BackgroundNon-protein-coding regions of eukaryotic genomes remain poorly understood. Diversity studies, comparative genomics and biochemical outputs of genomic sites can be indicators of functional elements, but none produce fine-scale genome-wide descriptions of all functional elements.\n\nResultsTowards the generation of a comprehensive description of functional elements in the haploid Schizosaccharomyces pombe genome, we generated transposon mutagenesis libraries to a density of one insertion per 13 nucleotides of the genome. We applied a five-state hidden Markov model (HMM) to characterise insertion-depleted regions at nucleotide-level resolution. HMM-defined functional constraint was consistent with genetic diversity, comparative genomics, gene-expression data and genome annotation.\n\nConclusionsWe infer that transposon insertions lead to fitness consequences in 90% of the genome, including 80% of the non-protein-coding regions, reflecting the presence of numerous non-coding elements in this compact genome that have functional roles. Display of this data in genome browsers provides fine-scale views of structure-function relationships within specific genes.

genomics

SKA: Split Kmer Analysis Toolkit for Bacterial Genomic Epidemiology

Genome sequencing is revolutionising infectious disease epidemiology, providing a huge step forward in sensitivity and specificity over more traditional molecular typing techniques. However, the complexity of genome data often means that its analysis and interpretation requires high-performance compute infrastructure and dedicated bioinformatics support. Furthermore, current methods have limitations that can differ between analyses and are often opaque to the user, and their reliance on multiple external dependencies makes reproducibility difficult. Here I introduce SKA, a toolkit for analysis of genome sequence data from closely-related, small, haploid genomes. SKA uses split kmers to rapidly identify variation between genome sequences, making it possible to analyse hundreds of genomes on a standard home computer. Tests on publicly available simulated and real-life data show that SKA is both faster and more efficient than the gold standard methods used today while retaining similar levels of accuracy for epidemiological purposes. SKA can take raw read data or genome assemblies as input and calculate pairwise distances, create single linkage clusters and align genomes to a reference genome or using a reference-free approach. SKA requires few decisions to be made by the user, which, along with its computational efficiency, allows genome analysis to become accessible to those with only basic bioinformatics training. The limitations of SKA are also far more transparent than for current approaches, and future improvements to mitigate these limitations are possible. Overall, SKA is a powerful addition to the armoury of the genomic epidemiologist. SKA source code is available from Github (https://github.com/simonrharris/SKA).

genomics

No evidence that sex and transposable elements drive genome size variation in evening primroses

Genome size varies dramatically across species, but despite an abundance of attention there is little agreement on the relative contributions of selective and neutral processes in governing this variation. The rate of sexual reproduction can potentially play an important role in genome size evolution because of its effect on the efficacy of selection and transmission of transposable elements. Here, we used a phylogenetic comparative approach and whole genome sequencing to investigate the contribution of sex and transposable element content to genome size variation in the evening primrose (Oenothera) genus. We determined genome size using flow cytometry from 30 Oenothera species of varying reproductive system and find that variation in sexual/asexual reproduction cannot explain the almost two-fold variation in genome size. Moreover, using whole genome sequences of three species of varying genome sizes and reproductive system, we found that genome size was not associated with transposable element abundance; instead the larger genomes had a higher abundance of simple sequence repeats. Although it has long been clear that sexual reproduction may affect various aspects of genome evolution in general and transposable element evolution in particular, it does not appear to have played a major role in the evening primroses.

Evolutionary Biology

gmos: Rapid detection of genome mosaicism over short evolutionary distances

Prokaryotic and viral genomes are often altered by recombination and horizontal gene transfer. The existing methods for detecting recombination are primarily aimed at viral genomes or sets of loci, since the expensive computation of underlying statistical models often hinders the comparison of complete prokaryotic genomes. As an alternative, alignment-free solutions are more efficient, but cannot map (align) a query to subject genomes. To address this problem, we have developed gmos (Genome MOsaic Structure), a new program that determines the mosaic structure of query genomes when compared to a set of closely related subject genomes. The program first computes local alignments between query and subject genomes and then reconstructs the query mosaic structure by choosing the best local alignment for each query region. To accomplish the analysis quickly, the program mostly relies on pairwise alignments and constructs multiple sequence alignments over short overlapping subject regions only when necessary. This fine-tuned implementation achieves an efficiency comparable to an alignment-free tool. The program performs well for simulated and real data sets of closely related genomes and can be used for fast recombination detection; for instance, when a new prokaryotic pathogen is discovered. As an example, gmos was used to detect genome mosaicism in a pathogenic Enterococcus faecium strain compared to seven closely related genomes. The analysis took less than two minutes on a single 2.1 GHz processor. The output is available in fasta format and can be visualized using an accessory program, gmosDraw (freely available with gmos).

Bioinformatics

Improved assemblies and comparison of two ancient Yersinia pestis genomes

Yersinia pestis is the causative agent of the bubonic plague, a disease responsible for several dramatic historical pandemics. Progress in ancient DNA (aDNA) sequencing rendered possible the sequencing of whole genomes of important human pathogens, including the ancient Yersinia pestis strains responsible for outbreaks of the bubonic plague in London in the 14th century and in Marseille in the 18th century among others. However, aDNA sequencing data are still characterized by short reads and non-uniform coverage, so assembling ancient pathogen genomes remains challenging and prevents in many cases a detailed study of genome rearrangements. It has recently been shown that comparative scaffolding approaches can improve the assembly of ancient Yersinia pestis genomes at a chromosome level. In the present work, we address the last step of genome assembly, the gap-filling stage. We describe an optimization-based method AGapEs (Ancestral Gap Estimation) to fill in inter-contig gaps using a combination of a template obtained from related extant genomes and aDNA reads. We show how this approach can be used to refine comparative scaffolding by selecting contig adjacencies supported by a mix of unassembled aDNA reads and comparative signal. We apply our method to two data sets from the London and Marseilles outbreaks of the bubonic plague. We obtain highly improved genome assemblies for both the London strain and Marseille strain genomes, comprised of respectively five and six scaffolds, with 95% of the assemblies supported by ancient reads. We analyze the genome evolution between both ancient genomes in terms of genome rearrangements, and observe a high level of synteny conservation between these two strains.

Bioinformatics

Detection and characterization of low and high genome coverage regions using an efficient running median and a double threshold approach.

MotivationNext Generation Sequencing (NGS) provides researchers with powerful tools to investigate both prokaryotic and eukaryotic genetics. An accurate assessment of reads mapped to a specific genome consists of inspecting the genome coverage as number of reads mapped to a specific genome location. Most current methods use the average of the genome coverage (sequencing depth) to summarize the overall coverage. This metric quickly assess the sequencing quality but ignores valuable biological information like the presence of repetitive regions or deleted genes. The detection of such information may be challenging due to a wide spectrum of heterogeneous coverage regions, a mixture of underlying models or the presence of a non-constant trend along the genome. Using robust statistics to systematically identify genomic regions with unusual coverage is needed to characterize these regions more precisely.\n\nResultsWe implemented an efficient running median algorithm to estimate the genome coverage trend. The distribution of the normalized genome coverage is then estimated using a Gaussian mixture model. A z-score statistics is then assigned to each base position and used to separate the central distribution from the regions of interest (ROI) (i.e., under and over-covered regions). Finally, a double threshold mechanism is used to cluster the genomic ROIs. HTML reports provide a summary with interactive visual representations of the genomic ROIs.\n\nAvailabilityAn implementation of the genome coverage characterization is available within the Sequana project. The standalone application is called sequana_coverage. The source code is available on GitHub (http://github.com/sequana/sequana), and documentation on ReadTheDocs (http://sequana.readtheodcs.org). An example of HTML report is provided on http://sequana.github.io.\n\nContactdimitri.desvillechabrol@pasteur.fr, thomas.cokelaer@pasteur.fr

bioinformatics

The 3D genome organization of Drosophila melanogaster through data integration

Genome structures are dynamic and non-randomly organized in the nucleus of higher eukaryotes. To maximize the accuracy and coverage of 3D genome structural models, it is important to integrate all available sources of experimental information about a genomes organization. It remains a major challenge to integrate such data from various complementary experimental methods. Here, we present an approach for data integration to determine a population of complete 3D genome structures that are statistically consistent with data from both genome-wide chromosome conformation capture (Hi-C) and lamina-DamID experiments. Our structures resolve the genome at the resolution of topological domains, and reproduce simultaneously both sets of experimental data. Importantly, this framework allows for structural heterogeneity between cells, and hence accounts for the expected plasticity of genome structures. As a case study we choose Drosophila melanogaster embryonic cells, for which both data types are available. Our 3D geome structures have strong predictive power for structural features not directly visible in the initial data sets, and reproduce experimental hallmarks of the D. melanogaster genome organization from independent and our own imaging experiments. Also they reveal a number of new insights about the genome organization and its functional relevance, including the preferred locations of heterochromatic satellites of differnet chromosomes, and observations about homologous pairing that cannot be directly observed in the original Hi-C or lamina-DamID data. To our knowledge our approach is the first that allows systematic integration of Hi-C and lamina DamID data for complete 3D genome structure calculation, while also explicitly considering genome structural variability.

bioinformatics

GAPPadder: A Sensitive Approach for Closing Gaps on Draft Genomes with Short Sequence Reads

BackgroundClosing gaps in draft genomes is an important post processing step in genome assembly. It leads to more complete genomes, which benefits downstream genome analysis such as annotation and genotyping. Several tools have been developed for gap closing. However, these tools dont fully utilize the information contained in the sequence data. For example, while it is known that many gaps are caused by genomic repeats, existing tools often ignore many sequence reads that originate from a repeat-related gap.\n\nResultsIn this paper, we propose a new approach called GAPPadder for gap closing. The main advantage of GAPPadder is that it uses more information in sequence data for gap closing. In particular, GAPPadder finds and uses reads that originate from repeate-related gaps. We show that these repeat-associated reads are useful for gap closing, even though they are ignored by all existing tools. Other main features of GAPPadder include utilizing the information in sequence reads with different insert sizes and performing two-stage local assembly of gap sequences. We compare GAPPadder with GapCloser, GapFiller and Sealer on one bacterial genome, human chromosome 14 and the human whole genome with paired-end and mate-paired reads with both short and long insert sizes. Empirical results show that GAPPadder can close more gaps than these existing tools. Besides closing gaps on draft genomes assembled only from short sequence reads, GAPPadder can also be used to close gaps for draft genomes assembled with long reads. We show GAPPadder can close gaps on the bed bug genome and the Asian sea bass genome that are assembled partially and fully with long reads respectively. We also show GAPPadder is efficient in both time and memory usage. The software tool, GAPPadder, is available for download at https://github.com/Reedwarbler/GAPPadder.

bioinformatics

Completing bacterial genome assemblies with multiplex MinION sequencing

Illumina sequencing platforms have enabled widespread bacterial whole genome sequencing. While Illumina data is appropriate for many analyses, its short read length limits its ability to resolve genomic structure. This has major implications for tracking the spread of mobile genetic elements, including those which carry antimicrobial resistance determinants. Fully resolving a bacterial genome requires long-read sequencing such as those generated by Oxford Nanopore Technologies (ONT) platforms. Here we describe our use of the ONT MinION to sequence 12 isolates of Klebsiella pneumoniae on a single flow cell. We assembled each genome using a combination of ONT reads and previously available Illumina reads, and little to no manual intervention was needed to achieve fully resolved assemblies using the Unicycler hybrid assembler. Assembling only ONT reads with Canu was less effective, resulting in fewer resolved genomes and higher error rates even following error correction with Nanopolish. We demonstrate that multiplexed ONT sequencing is a valuable tool for high-throughput bacterial genome finishing. Specifically, we advocate the use of Illumina sequencing as a first analysis step, followed by ONT reads as needed to resolve genomic structure.\n\nData summaryO_LISequence read files for all 12 isolates have been deposited in SRA, accessible through these NCBI BioSample accession numbers: SAMEA3357010, SAMEA3357043, SAMN07211279, SAMN07211280, SAMEA3357223, SAMEA3357193, SAMEA3357346, SAMEA3357374, SAMEA3357320, SAMN07211281, SAMN07211282, SAMEA3357405.\nC_LIO_LIA full list of SRA run accession numbers (both Illumina reads and ONT reads) for these samples are available in Table S1.\nC_LIO_LIAssemblies and sequencing reads corresponding to each stage of processing and analysis are provided in the following figshare project: https://figshare.com/projects/Completing_bacterial_genome_assemblies_with_multiplex_MinION_sequencing/23068\nC_LIO_LISource code is provided in the following public GitHub repositories: https://github.com/rrwick/Bacterial-genome-assemblies-with-multiplex-MinION-sequencing https://github.com/rrwick/Porechop https://github.com/rrwick/Fast5-to-Fastq\nC_LI\n\nImpact StatementLike many research and public health laboratories, we frequently perform large-scale bacterial comparative genomics studies using Illumina sequencing, which assays gene content and provides the high-confidence variant calls needed for phylogenomics and transmission studies. However, problems often arise with resolving genome assemblies, particularly around regions that matter most to our research, such as mobile genetic elements encoding antibiotic resistance or virulence genes. These complexities can often be resolved by long sequence reads generated with PacBio or Oxford Nanopore Technologies (ONT) platforms. While effective, this has proven difficult to scale, due to the relatively high costs of generating long reads and the manual intervention required for assembly. Here we demonstrate the use of barcoded ONT libraries sequenced in multiplex on a single ONT MinION flow cell, coupled with hybrid assembly using Unicycler, to resolve 12 large bacterial genomes. Minor manual intervention was required to fully resolve small plasmids in five isolates, which we found to be underrepresented in ONT data. Cost per sample for the ONT sequencing was equivalent to Illumina sequencing, and there is potential for significant savings by multiplexing more samples on the ONT run. This approach paves the way for high-throughput and cost-effective generation of completely resolved bacterial genomes to become widely accessible.

bioinformatics

Genomic and proteomic analysis of Human herpesvirus 6 reveals distinct clustering of acute versus inherited forms and reannotation of reference strain

Human herpesvirus-6A and -6B (HHV-6) are betaherpesviruses that reach >90% seroprevalence in the adult population. Unique among human herpesviruses, HHV-6 can integrate into the subtelomeric regions of human chromosomes; when this occurs in germ line cells it causes a condition called inherited chromosomally integrated HHV-6 (iciHHV-6). To date, only two complete genomes are available for HHV-6B. Using a custom capture panel for HHV-6B, we report near-complete genomes from 61 isolates of HHV-6B from active infections (20 from Japan, 35 from New York state, and 6 from Uganda), and 64 strains of iciHHV-6B (mostly from North America). We also report partial genome sequences from 10 strains of iciHHV-6A. Although the overall sequence diversity of HHV-6 is limited relative to other human herpesviruses, our sequencing identified geographical clustering of HHV-6B sequences from active infections, as well as evidence of recombination among HHV-6B strains. One strain of active HHV-6B was more divergent than any other HHV-6B previously sequenced. In contrast to the active infections, sequences from iciHHV-6 cases showed reduced sequence diversity. Strikingly, multiple iciHHV-6B sequences from unrelated individuals were found to be completely identical, consistent with a founder effect. However, several iciHHV-6B strains intermingled with strains from active pediatric infection, consistent with the hypothesis that intermittent de novo integration into host germline cells can occur during active infection Comparative genomic analysis of the newly sequenced strains revealed numerous instances where conflicting annotations between the two existing reference genomes could be resolved. Combining these findings with transcriptome sequencing and shotgun proteomics, we reannotated the HHV-6B genome and found multiple instances of novel splicing and genes that hitherto had gone unannotated. The results presented here constitute a significant genomic resource for future studies on the detection, diversity, and control of HHV-6.\n\nAuthor SummaryHHV-6 is a ubiquitous large DNA virus that is the most common cause of febrile seizures and reactivates in allogeneic stem cell patients. It also has the unique ability among human herpesviruses to be integrated into the genome of every cell via integration in the germ line, a condition called inherited chromosomally integrated (ici)HHV-6, which affects approximately 1% of the population. To date, very little is known about the comparative genomics of HHV-6. We sequenced 61 isolates of HHV-6B from active infections, 64 strains of iciHHV-6B, and 10 strains of iciHHV-6A. We found geographic clustering of HHV-6B strains from active infections. In contrast, iciHHV-6B had reduced sequence diversity, with many identical sequences of iciHHV-6 found in individuals not known to share recent common ancestry, consistent with a founder effect from a remote common ancestor with iciHHV-6. We also combined our genomic analysis with transcriptome sequencing and shotgun proteomics to correct previous misannotations of the HHV-6 genome.

microbiology

A Fast Adaptive Algorithm for Computing Whole-Genome Homology Maps

MotivationWhole-genome alignment is an important problem in genomics for comparing different species, mapping draft assemblies to reference genomes, and identifying repeats. However, for large plant and animal genomes, this task remains compute and memory intensive.\n\nResultsWe introduce an approximate algorithm for computing local alignment boundaries between long DNA sequences. Given a minimum alignment length and an identity threshold, our algorithm computes the desired alignment boundaries and identity estimates using kmer-based statistics, and maintains sufficient probabilistic guarantees on the output sensitivity. Further, to prioritize higher scoring alignment intervals, we develop a plane-sweep based filtering technique which is theoretically optimal and practically efficient. Implementation of these ideas resulted in a fast and accurate assembly-to-genome and genome-to-genome mapper. As a result, we were able to map an error-corrected whole-genome NA12878 human assembly to the hg38 human reference genome in about one minute total execution time and < 4 GB memory using 8 CPU threads, achieving more than an order of magnitude improvement in both runtime and memory over competing methods. Recall accuracy of computed alignment boundaries was consistently found to be > 97% on multiple datasets. Finally, we performed a sensitive self-alignment of the human genome to compute all duplications of length [&ge;] 1 Kbp and [&ge;] 90% identity. The reported output achieves good recall and covers 5% more bases than the current UCSC genome browsers segmental duplication annotation.\n\nAvailabilityhttps://github.com/marbl/MashMap\n\nContactadam.phillippy@nih.gov, aluru@cc.gatech.edu

bioinformatics

FusoPortal: An interactive repository of hybrid MinION sequenced Fusobacterium genomes improves gene identification and characterization

Here we present FusoPortal, an interactive repository of Fusobacterium genomes that were sequenced using a hybrid MinION long-read sequencing pipeline, followed by assembly and annotation using a diverse portfolio of predominantly open-source software. Significant efforts were made to provide genomic and bioinformatic data as downloadable files, including raw sequencing reads, genome maps, gene annotations, protein functional analysis and classifications, and a custom BLAST server for FusoPortal genomes. FusoPortal has been initiated with eight complete genomes, of which seven were previously only drafts that varied from 24-67 contigs. We showcase that genomes in FusoPortal provide accurate open reading frame annotations, and have corrected a number of large genes (>3 kb) that were previously misannotated due to contig boundaries. In summary, FusoPortal (http://fusoportal.org) is the first database of MinION sequenced and completely assembled Fusobacterium genomes, and this central Fusobacterium genomic and bioinformatic resource will aid the scientific community in developing a deeper understanding of how this human pathogen contributes to an array of diseases including periodontitis and colorectal cancer.\n\nImportanceIn this study, we report a hybrid MinION whole genome sequencing pipeline, and describe the genomic characteristics of the first eight strains deposited in the FusoPortal database. This collection of highly accurate and complete genomes drastically improves upon previous multi-contig assemblies by correcting and newly identifying a significant number of open reading frames. We believe this resource will result in the discovery of proteins and molecular mechanisms used by an oral pathogen, with the potential to further our understanding of how F. nucleatum contributes to a repertoire of diseases including periodontitis, pre-term birth, and colorectal cancer

microbiology

Genome plasticity, a key factor of evolution in prokaryotes

In prokaryotic genomes, the number of genes that belong to distinct functional classes shows apparent universal scaling with the total number of genes [1-5] (Fig. 1). This scaling can be approximated with a power law, where the scaling power can be sublinear, near-linear or super-linear. Scaling laws are robust under various statistical tests [4], across different databases and for different gene classifications [1-5]. Several models aimed at explaining the observed scaling laws have been proposed, primarily, based on the specifics of the respective biological functions [1, 5-8]. However, a coherent theory to explain the emergence of scaling within the framework of population genetics is lacking. We employ a simple mathematical model for prokaryotic genome evolution [9] which, together with the analysis of 34 clusters of closely related microbial genomes [10], allows us to identify the underlying forces that dictate genome content evolution. In addition to the scaling of the number of genes in different functional classes, we explore gene contents divergence to characterize the evolutionary processes acting upon genomes [11]. We find that evolution of the gene content is dominated by two factors that are specific to a functional class, namely, selection landscape and genome plasticity. Selection landscape quantifies the fitness cost that is associated with deletion of a gene in a given functional class or the advantage of successful incorporation of an additional gene. Genome plasticity, that can be considered a measure of evolvability, reflects both the availability of the genes of a given functional class in the external gene pool that is accessible to the evolving microbial population, and the ability of microbial genomes to accommodate these genes. The selection landscape determines the gene loss rate, and genome plasticity is the principal determinant of the gene gain rate.\n\nO_FIG O_LINKSMALLFIG WIDTH=197 HEIGHT=200 SRC=\"FIGDIR/small/357400_fig1.gif\" ALT=\"Figure 1\">\nView larger version (39K):\norg.highwire.dtl.DTLVardef@6df3e2org.highwire.dtl.DTLVardef@a69e8dorg.highwire.dtl.DTLVardef@f36a80org.highwire.dtl.DTLVardef@d519c9_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Scaling laws for all functional classes of the COGs. The number of genes in a given COG category is plotted against the total number of genes. Each point represents one genome from the analyzed set of 1490 genomes. The scaling is fitted to a power law which is indicated by a solid red line. The fitted scaling exponent is indicated in parentheses.\n\nC_FIG

microbiology

RT States: systematic annotation of the human genome using cell type-specific replication timing programs

The replication timing (RT) program has been linked to many key biological processes including cell fate commitment, 3D chromatin organization and transcription regulation. Significant technology progress now allows to characterize the RT program in the entire human genome in a high-throughput and high-resolution fashion. These experiments suggest that RT changes dynamically during development in coordination with gene activity. Since RT is such a fundamental biological process, we believe that an effective quantitative profile of the local RT program from a diverse set of cell types in various developmental stages and lineages can provide crucial biological insights for a genomic locus. In the present study, we explored recurrent and spatially coherent combinatorial profiles from 42 RT programs collected from multiple lineages at diverse differentiation states. We found that a Hidden Markov Model with 15 hidden states provide a good model to describe these genome-wide RT profiling data. Each of the hidden state represents a unique combination of RT profiles across different cell types which we refer to as \"RT states\". To understand the biological properties of these RT states, we inspected their relationship with chromatin states, gene expression, functional annotation and 3D chromosomal organization. We found that the newly defined RT states possess interesting genome-wide functional properties that add complementary information to the existing annotation of the human genome.\n\nAUTHOR SUMMARYThe replication timing (RT) program is an important cellular mechanism and has been linked to many key biological processes including cell fate commitment, 3D chromatin organization and transcription regulation. Significant technology progress now allows us to characterize the RT program in the entire human genome. Results from these experiments suggest that RT changes dynamically across different developmental stages. Since RT is such a fundamental biological process, we believe that the local RT program from a diverse set of cell types in various developmental stages can provide crucial biological insights for a genomic locus. In the present study, we explored combinatorial profiles from 42 RT programs collected from multiple lineages at diverse differentiation states. We developed a statistical model consist of 15 \"RT states\" to describe these genome-wide RT profiling data. To understand the biological properties of these RT states, we inspected the relationship between RT states and other types of functional annotations of the genome. We found that the newly defined RT states possess interesting genome-wide functional properties that add complementary information to the existing annotation of the human genome.

bioinformatics