bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

An open library of human kinase domain constructs for automated bacterial expression

Kinases play a critical role in many cellular signaling pathways and are dysregulated in a number of diseases, such as cancer, diabetes, and neurodegeneration. Since the FDA approval of imatinib in 2001, therapeutics targeting kinases now account for roughly 50% of current cancer drug discovery efforts. The ability to explore human kinase biochemistry, biophysics, and structural biology in the laboratory is essential to making rapid progress in understanding kinase regulation, designing selective inhibitors, and studying the emergence of drug resistance. While insect and mammalian expression systems are frequently used for the expression of human kinases, bacterial expression systems are superior in terms of simplicity and cost-effectiveness but have historically struggled with human kinase expression. Following the discovery that phosphatase coexpression could produce high yields of Src and Abl kinase domains in bacterial expression systems, we have generated a library of 52 His-tagged human kinase domain constructs that express above 2 {micro}g/mL culture in a simple automated bacterial expression system utilizing phosphatase coexpression (YopH for Tyr kinases, Lambda for Ser/Thr kinases). Here, we report a structural bioinformatics approach to identify kinase domain constructs previously expressed in bacteria likely to express well in a simple high-throughput protocol, experiments demonstrating our simple construct selection strategy selects constructs with good expression yields in a test of 84 potential kinase domain boundaries for Abl, and yields from a high-throughput expression screen of 96 human kinase constructs. Using a fluorescence-based thermostability assay and a fluorescent ATP-competitive inhibitor, we show that the highest-expressing kinases are folded and have well-formed ATP binding sites. We also demonstrate how the resulting expressing constructs can be used for the biophysical and biochemical study of clinical mutations by engineering a panel of 48 Src mutations and 46 Abl mutations via single-primer mutagenesis and screening the resulting library for expression yields. The wild-type kinase construct library is available publicly via Addgene, and should prove to be of high utility for experiments focused on drug discovery and the emergence of drug resistance.

Biochemistry↗

DNA from dust: comparative genomics of large DNA viruses in field surveillance samples

Mareks disease (MD) is a lymphoproliferative disease of chickens caused by airborne gallid herpesvirus type 2 (GaHV-2, aka MDV-1). Mature virions are formed in the feather follicle epithelium cells of infected chickens from which the virus is shed as fine particles of skin and feather debris, or poultry dust. Poultry dust is the major source of virus transmission between birds in agricultural settings. Despite both clinical and laboratory data that show increased virulence in field isolates of MDV-1 over the last 40 years, we do not yet understand the genetic basis of MDV-1 pathogenicity. Our present knowledge on genome-wide variation in the MDV-1 genome comes exclusively from laboratory-grown isolates. MDV-1 isolates tend to lose virulence with increasing passage number in vitro, raising concerns about their ability to accurately reflect virus in the field. The ability to rapidly and directly sequence field isolates of MDV-1 is critical to understanding the genetic basis of rising virulence in circulating wild strains. Here we present the first complete genomes of uncultured, field-isolated MDV-1. These five consensus genomes were derived directly from poultry dust or single chicken feather follicles without passage in cell culture. These sources represent the shed material that is transmitted to new hosts, vs. the virus produced by a point source in one animal. We developed a new procedure to extract and enrich viral DNA, while reducing host and environmental contamination. DNA was sequenced using Illumina MiSeq high-throughput approaches and processed through a recently described bioinformatics workflow for de novo assembly and curation of herpesvirus genomes. We comprehensively compared these genomes to one another and also to previously described MDV-1 genomes. The field-isolated genomes had remarkably high DNA identity when compared to one another, with few variant proteins between them. In an analysis of genetic distance, the five new field genomes grouped separately from all previously described genomes. Each consensus genome was also assessed to determine the level of polymorphisms within each sample, which revealed that MDV-1 exists in the wild as a polymorphic population. By tracking a new polymorphic locus in ICP4 over time, we found that MDV-1 genomes can evolve in short period of time. Together these approaches advance our ability to assess MDV-1 variation within and between hosts, over time, and during adaptation to changing conditions.

Microbiology↗

Assemblytics: a web analytics tool for the detection of assembly-based variants

SummaryAssemblytics is a web app for detecting and analyzing structural variants from a de novo genome assembly aligned to a reference genome. It incorporates a unique anchor filtering approach to increase robustness to repetitive elements, and identifies six classes of variants based on their distinct alignment signatures. Assemblytics can be applied both to comparing aberrant genomes, such as human cancers, to a reference, or to identify differences between related species. Multiple interactive visualizations enable in-depth explorations of the genomic distributions of variants.\n\nAvailability and Implementationhttp://qb.cshl.edu/assemblytics, https://github.com/marianattestad/assemblytics\n\nContact: mnattest@cshl.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Genomics↗

Functional metagenomics reveals novel β-galactosidases not predictable from gene sequences

A soil metagenomic library carried in pJC8 (an IncP cosmid) was used for functional complementation for {beta}-galactosidase activity in both -Proteobacteria (Sinorhizobium meliloti) and{gamma} -Proteobacteria (Escherichia coli). One {beta}-galactosidase, encoded by overlapping clones selected in both hosts, was identified as a member of glycoside hydrolase family 2. ORFs obviously encoding possible {beta}-galactosidases were not identified in 19 other clones that were only able to complement S. meliloti. Based on low sequence similarity to known glycoside hydrolases but not {beta}-galactosidases, three ORFs were examined further. Biochemical analysis confirmed that all encoded {beta}-galactosidase activity. Bioinformatic and structural modeling implied that Lac161_ORF10 protein represented a novel enzyme family with a five-bladed propeller glycoside hydrolase domain.

Microbiology↗

Characterization of sterol synthesis in bacteria

Sterols are essential components of eukaryotic cells whose biosynthesis and function in eukaryotes has been studied extensively. Sterols are also recognized as the diagenetic precursors of steranes preserved in sedimentary rocks where they can function as geological proxies for eukaryotic organisms and/or aerobic metabolisms and environments. However, production of these lipids is not restricted to the eukaryotic domain as a few bacterial species also synthesize sterols. Phylogenomic studies have identified genes encoding homologs of sterol biosynthesis proteins in the genomes of several additional species, indicating that sterol production may be more widespread in the bacterial domain than previously thought. Although the occurrence of sterol synthesis genes in a genome indicates the potential for sterol production, it provides neither conclusive evidence of sterol synthesis nor information about the composition and abundance of basic and modified sterols that are actually being produced. Here, we coupled bioinformatics with lipid analyses to investigate the scope of bacterial sterol production. We identified oxidosqualene cyclase (Osc), which catalyzes the initial cyclization of oxidosqualene to the basic sterol structure, in 34 bacterial genomes from 5 phyla (Bacteroidetes, Cyanobacteria, Planctomycetes, Proteobacteria and Verrucomicrobia) and in 176 metagenomes. Our data indicate that bacterial sterol synthesis likely occurs in diverse organisms and environments and also provides evidence that there are as yet uncultured groups of bacterial sterol producers. Phylogenetic analysis of bacterial and eukaryotic Osc sequences revealed two potential lineages of the sterol pathway in bacteria indicating a complex evolutionary history of sterol synthesis in this domain. We characterized the lipids produced by Osc-containing bacteria and found that we could generally predict the ability to synthesize sterols. However, predicting the final modified sterol based on our current knowledge of bacterial sterol synthesis was difficult. Some bacteria produced demethylated and saturated sterol products even though they lacked homologs of the eukaryotic proteins required for these modifications emphasizing that several aspects of bacterial sterol synthesis are still completely unknown. It is possible that bacteria have evolved distinct proteins for catalyzing sterol modifications and this could have significant implications for our understanding of the evolutionary history of this ancient biosynthetic pathway.

Microbiology↗

Analytical considerations for comparative transcriptomics of wild organisms.

Comparative transcriptomics can now be conducted on organisms in natural settings, which has greatly enhanced understanding of genome-environment interactions. However, important data handling and quality control challenges remain, particularly when working with non-model species outside of a controlled laboratory environment. Here, we demonstrate the utility and potential pitfalls of comparative transcriptomics of wild organisms, with an example from three cyprinid fish species (Teleostei:Cypriniformes). We present computational solutions for processing, annotating and summarizing comparative transcriptome data for assessing genome-environment interactions across species. The resulting bioinformatics pipeline addresses the following points: (1) the potential importance of \"essential genes\", (2) the influence of microbiomes and other exogenous DNA, (3) potentially novel, species-specific genes, and (4) genomic rearrangements (e.g., whole genome duplication). Quantitative consideration of these points contributes to a firmer foundation for future comparative work across distantly related taxa for a variety of sub-disciplines, including stress and immune response, community ecology, ecotoxicology, and climate change.

Genomics↗

Optimizing multiplex CRISPR/Cas9-based genome editing for wheat

BackgroundCRISPR/Cas9-based genome editing holds great promise to accelerate the development of new crop varieties by providing a powerful tool to modify the genomic regions controlling major agronomic traits. To diversify the set of tools available for wheat genome engineering, we have established a tRNA-based multiplex gene editing strategy for hexaploid wheat.\n\nResultsThe functionality of the various CRISPR/Cas9 components was assessed using the transient expression in the wheat protoplasts followed by next-generation sequencing (NGS) of the targeted genomic regions. The efficiency of wheat codon-optimized Cas9 for targeted gene editing in wheat was validated. Multiple single guide RNAs (gRNAs) were evaluated for the ability to edit the homoeologous copies of four genes affecting some important agronomic traits in wheat. Low correspondence was found between the gRNA efficiency predicted bioinformatically and that assessed in the transient expression assay. A multiplex gene editing construct with several gRNA-tRNA units under the control of a single promoter for the RNA polymerase III generated indels at the targets sites with the efficiency comparable to that obtained for a single gRNA construct.\n\nConclusionsBy integrating the protoplast transformation assay with multiplexed NGS, it is possible to perform fast functional screens for a large number of gRNAs and to optimize constructs for effective editing of multiple independent targets in the wheat genome. The multiplexing capacity of the tandemly arrayed tRNA-gRNA construct is well suited for the simultaneous editing of the redundant gene copies in the allopolyploid genomes or genomic regions beneficially affecting multiple agronomic traits. A polycistronic gene construct that can be quickly assembled using the Golden Gate reaction along with the wheat codon optimized Cas9 will further expand the set of tools available for engineering the wheat genome.

Genomics↗

Scalable Design of Paired CRISPR Guide RNAs for Genomic Deletion

Using CRISPR/Cas9, diverse genomic elements may be studied in their endogenous context. Pairs of single guide RNAs (sgRNAs) are used to delete regulatory elements and small RNA genes, while longer RNAs can be silenced through promoter deletion. We here present CRISPETa, a bioinformatic pipeline for flexible and scalable paired sgRNA design based on an empirical scoring model. Multiple sgRNA pairs are returned for each target. Any number of targets can be analyzed in parallel, making CRISPETa equally appropriate for studies of individual elements, or complex library screens. Fast run-times are achieved using a precomputed off-target database. sgRNA pair designs are output in a convenient format for visualisation and oligonucleotide ordering. We present a series of pre-designed, high-coverage library designs for entire classes of non-coding elements in human, mouse, zebrafish, Drosophila and C. elegans. Using an improved version of the DECKO deletion vector, together with a quantitative deletion assay, we test CRISPETa designs by deleting an enhancer and exonic fragment of the MALAT1 oncogene. These achieve efficiencies of [≥]50%, resulting in production of mutant RNA. CRISPETa will be useful for researchers seeking to harness CRISPR for targeted genomic deletion, in a variety of model organisms, from single-target to high-throughput scales.

Genomics↗

Suitability of different mapping algorithms for genome-wide polymorphism scans with Pool-Seq data

The cost-effectiveness of sequencing pools of individuals (Pool-Seq) provides the basis for the popularity and wide-spread use of this method for many research questions, ranging from unravelling the genetic basis of complex traits to the clonal evolution of cancer cells. Because the accuracy of Pool-Seq could be affected by many potential sources of error, several studies determined, for example, the influence of the sequencing technology, the library preparation protocol, and mapping parameters. Nevertheless, the impact of the mapping tools has not yet been evaluated. Using simulated and real Pool-Seq data, we demonstrate a substantial impact of the mapping tools leading to characteristic false positives in genome-wide scans. The problem of false positives was particularly pronounced when data with different read lengths and insert sizes were compared. Out of 14 evaluated algorithms novoalign, bwa mem and clc4 are most suitable for mapping Pool-Seq data. Nevertheless, no single algorithm is sufficient for avoiding all false positives. We show that the intersection of the results of two mapping algorithms provides a simple, yet effective strategy to eliminate false positives. We propose that the implementation of a consistent Pool-seq bioinformatics pipeline building on the recommendations of this study can substantially increase the reliability of Pool-Seq results, in particular when libraries generated with different protocols are being compared.

Genomics↗

Phages rarely encode antibiotic resistance genes: a cautionary tale for virome analysis

Antibiotic resistance genes (ARG) are pervasive in gut microbiota, but it remains unclear how often ARG are transferred, particularly to pathogens. Traditionally, ARG spread is attributed to horizontal transfer mediated either by DNA transformation, bacterial conjugation or generalized transduction. However, recent viral metagenome (virome) analyses suggest that ARG are frequently carried by phages, which is inconsistent with the traditional view that phage genomes rarely encode ARG. Here we used exploratory and conservative bioinformatic strategies found in the literature to detect ARG in phage genomes, and experimentally assessed a subset of ARG predicted using exploratory thresholds. ARG abundances in 1,181 phage genomes were vastly over-estimated using exploratory thresholds (421 predicted vs 2 known), due to low similarities and matches to protein unrelated to antibiotic resistance. Consistent with this, 4 ARG predicted using exploratory thresholds were experimentally evaluated and failed to confer antibiotic resistance in Escherichia coli. Re-analysis of available human-or mouse-associated viromes for ARG and their genomic context suggested that bona fide ARG attributed to phages in viromes were previously over-estimated. These findings provide guidance for documentation of ARG in viromes, and re-assert that ARG are rarely encoded in phages.

Ecology↗

Comparing the Statistical Fate of Paralogous and Orthologous Sequences

Since several decades, sequence alignment is a widely used tool in bioinformatics. For instance, finding homologous sequences with known function in large databases is used to get insight into the function of non-annotated genomic regions. Very efficient tools, like BLAST have been developed to identify and rank possible homologous sequences. To estimate the significance of the homology, the ranking of alignment scores takes a background model for random sequences into account. Using this model one can estimate the probability to find two exactly matching subsequences by chance in two unrelated sequences. The corresponding probability for two homologous sequences is much higher allowing to identify them. Here we focus on the distribution of lengths of exact sequence matches in protein coding regions pairs of evolutionary distant genomes. We show that this distribution exhibits a power-law tail with exponent = --5. Developing a simple model of sequence evolution by substitutions and segmental duplications, we show analytically that paralogous and orthologous gene pairs contribute differently to this distribution. Our model explains the differences observed in the comparison of coding and non-coding parts of genomes, thus providing with a better understanding of statistical properties of genomic sequences and their evolution.

Evolutionary Biology↗

Steady at the wheel: conservative sex and the benefits of bacterial transformation

Many bacteria are highly sexual, but the reasons for their promiscuity remain obscure. Did bacterial sex evolve to maximize diversity and facilitate adaptation in a changing world, or does it instead help to retain the bacterial functions that work right now? In other words, is bacterial sex innovative or conservative? Our aim in this review is to integrate experimental, bioinformatic and theoretical studies to critically evaluate these alternatives, with a main focus on natural genetic transformation, the bacterial equivalent of eukaryotic sexual reproduction. First, we provide a general overview of several hypotheses that have been put forward to explain the evolution of transformation. Next, we synthesize a large body of evidence highlighting the numerous passive and active barriers to transformation that have evolved to protect bacteria from foreign DNA, thereby increasing the likelihood that transformation takes place among clonemates. Our critical review of the existing literature provides support for the view that bacterial transformation is maintained as a means of genomic conservation that provides direct benefits to both individual bacterial cells and to transformable bacterial populations. We examine the generality of this view across bacteria and contrast this explanation with the different evolutionary roles proposed to maintain sex in eukaryotes.

Evolutionary Biology↗

Persistent activation of interlinked Th2-airway epithelial gene networks in sputum-derived cells from aeroallergen-sensitized symptomatic atopic asthmatics

RationaleAtopic asthma is a persistent disease characterized by intermittent wheeze and progressive loss of lung function. The disease is thought to be driven primarily by chronic aeroallergen-induced Th2-associated airways inflammation. However, the vast majority of atopics do not develop asthma-related wheeze, despite ongoing exposure to aeroallergens to which they are strongly sensitized, indicating that additional pathogenic mechanism(s) operate in conjunction with Th2 immunity to drive asthma pathogenesis.\n\nObjectivesEmploy systems level analyses to identify inflammation-associated gene networks operative at baseline in sputum-derived RNA from house dust mite-sensitized (HDMs) subjects with/without wheezing history; identify networks characteristic of the ongoing asthmatic state. All subjects resided in the constitutively-HDMhigh Perth environment.\n\nMethodsGenome wide expression profiling by RNASeq followed by gene coexpression network analysis.\n\nMeasurements/ResultsHDMs-nonwheezers displayed baseline gene expression in sputum including IL-5, IL-13 and CCL17. HDMs-wheezers showed equivalent expression of these classical Th2-effector genes but their overall baseline sputum signatures were more complex, comprising hundreds of Th2-associated and epithelial-associated genes, networked into two separate coexpression modules. The first module was connected by the hubs EGFR, ERBB2, CDH1 and IL-13. The second module was associated with CDHR3, and contained genes that control mucociliary clearance.\n\nConclusionsOur findings provide new insight into the inflammatory mechanisms operative at baseline in the airway mucosal microenvironment in atopic asthmatics undergoing natural perennial aeroallergen exposure. The molecular mechanism(s) that determine susceptibility to asthma amongst these subjects involve interactions between Th2-and epithelial function-associated genes within a complex co-expression network, which is not operative in equivalently sensitized/exposed atopic non-asthmatics.\n\nFundingThis study was funded by the Asthma Foundation WA, the Department of Health WA, and the NHMRC. AB is funded by a BrightSpark Foundation McCusker Fellowship. GLH is a NHMRC Fellow. AG is supported by the McCusker Charitable Foundation Bioinformatics Centre. ACJ is a recipient of an Australian Postgraduate Award and a Top-Up Award from the University of Western Australia.

Immunology↗

Pherotype polymorphism in Streptococcus pneumoniae and its effects on population structure and recombination

Natural transformation in the Gram-positive pathogen Streptococcus pneumoniae occurs when cells become \"competent\", a state that is induced in response to high extracellular concentrations of a secreted peptide signal called CSP (Competence Stimulating Peptide) encoded by the comC locus. Two main CSP signal types (pherotypes) are known to dominate the pherotype diversity across strains. Using thousands of fully sequenced pneumococcal genomes, we confirm that pneumococcal populations are highly genetically structured and that there is significant variation among diverged populations in pherotype frequencies; most carry only a single pherotype. Moreover, we find that the relative frequencies of the two dominant pherotypes significantly vary within a small range across geographical sites. It has been variously proposed that pherotypes either promote genetic exchange among cells expressing the same pherotype, or conversely that they promote recombination between strains bearing different pherotypes. We distinguish these hypotheses using a bioinformatics approach by estimating recombination frequencies within and between pherotypes across 4,089 full genomes. Despite underlying population structure, we observe extensive recombination between populations; additionally, we found significantly higher rates of genetic exchange between strains expressing different pherotypes than among isolates carrying the same pherotype. Our results indicate that pherotypes do not restrict, and marginally facilitate, recombination between strains. Furthermore, our results suggest that the CSP balanced polymorphism does not causally underlie population differentiation. Therefore, when strains carrying different pherotypes encounter one another during co-colonization, genetic exchange can freely occur.

Evolutionary Biology↗

Evaluating Mendelian nephrotic syndrome genes for evidence of risk alleles or oligogenicity that explain heritability

BackgroundMore than 30 genes can harbor rare exonic variants sufficient to cause nephrotic syndrome (NS), and the number of genes implicated in monogenic NS continues to grow. However, outside the first year of life, the majority of affected patients, particularly in ancestrally mixed populations, do not have a known monogenic form of NS. Even in those children classified with a monogenic form of NS, there is phenotypic heterogeneity. Thus, we have only discovered a fraction of the heritability of NS - the underlying genetic factors contributing to phenotypic variation. Part of the \"missing heritability\" for NS has been posited to be explained by patients harboring coding variants across one or more previously implicated NS genes, insufficient to cause NS in a classical Mendelian manner, but that nonetheless impact protein function enough to cause disease. However, systematic evaluation in patients with NS for rare or low-frequency risk alleles within single genes, or in combination across genes (\"oligogenicity\"), has not been reported.\n\nObjectiveTo determine whether, as compared to a reference population, patients with NS have either a significantly increased burden of protein-altering variants (\"risk-alleles\"), or unique combination of them (\"oligogenicity\"), in a set of 21 genes implicated in Mendelian forms of NS.\n\nMethodsIn 303 patients with NS enrolled in the Nephrotic Syndrome Study Network (NEPTUNE), we performed targeted amplification paired with next-generation sequencing of 21 genes implicated in monogenic NS. We created a high-quality variant call set and compared it to a variant call set of the same genes in a reference population composed of 2535 individuals from Phase 3 of 1000 Genomes Project. We created both a \"stringent\" and \"relaxed\" pathogenicity filtering pipeline, applied them to both cohorts, and computed the (1) burden of variants in the entire gene set per cohort, (2) burden of variants in the entire gene set per individual, (3) burden of variants within a single gene per cohort, and (4) unique combinations of variants across two or more genes per cohort.\n\nResultsWith few exceptions when using the relaxed filter, and which are likely the result of confounding by population stratification, NS patients did not have significantly increased burden of variants in Mendelian NS genes in comparison to a reference cohort, nor was there any evidence of oligogenicity. This was true when using both the relaxed and stringent variant pathogenicity filter.\n\nConclusionIn our study, the burden or particular combinations of low-frequency or rare protein altering variants in previously implicated Mendelian NS genes cohort does not significantly differ between North American patients with NS and a reference population. Studies in larger independent cohorts or meta-analyses are needed to assess generalizability of our discoveries and also address whether there is in fact small but significant enrichment of risk alleles or oligogenicity in NS cases undetectable with this current sample size. It is still possible that rare protein altering variants in these genes, insufficient to cause Mendelian disease, still contribute to NS as risk alleles and/or via oligogenicity. However, we suggest that more accurate bioinformatic analyses and the incorporation of functional assays would be necessary to identify bona fide instances of this form of genetic architecture as a contributor to the heritability of NS.

Genetics↗

Nuclear pore-like structures in a compartmentalized bacterium

Planctomycetes are distinguished from other Bacteria by compartmentalization of cells via internal membranes, interpretation of which has been subject to recent debate regarding potential relations to Gram-negative cell structure. In our interpretation of the available data, the planctomycete Gemmata obscuriglobus contains a nuclear body compartment, and thus possesses a type of cell organization with parallels to the eukaryote nucleus. Here we show that pore-like structures occur in internal membranes of G.obscuriglobus and that they have elements structurally similar to eukaryote nuclear pores, including a basket, ring-spoke structure, and eight-fold rotational symmetry. Bioinformatic analysis of proteomic data reveals that some of the G. obscuriglobus proteins associated with pore-containing membranes possess structural domains found in eukaryote nuclear pore complexes. Moreover, immuno-gold labelling demonstrates localization of one such protein, containing a {beta}-propeller domain, specifically to the G. obscuriglobus pore-like structures. Finding bacterial pores within internal cell membranes and with structural similarities to eukaryote nuclear pore complexes raises the dual possibilities of either hitherto undetected homology or stunning evolutionary convergence.

Microbiology↗

SiLiCO: A Simulator of Long Read Sequencing in PacBio and Oxford Nanopore

SummaryLong read sequencing platforms, which include the widely used Pacific Biosciences (PacBio) platform and the emerging Oxford Nanopore platform, aim to produce sequence fragments in excess of 15-20 kilobases, and have proved advantageous in the identification of structural variants and easing genome assembly. However, long read sequencing remains relatively expensive and error prone, and failed sequencing runs represent a significant problem for genomics core facilities. To quantitatively assess the underlying mechanics of sequencing failure, it is essential to have highly reproducible and controllable reference data sets to which sequencing results can be compared. Here, we present SiLiCO, the first in silico simulation tool to generate standardized sequencing results from both of the leading long read sequencing platforms.\n\nAvailabilitySiLiCO is an open source package written in Python. It is freely available at https://www.github.com/ethanagbaker/SiLiCO under the GNU GPL 3.0 license.\n\nContact \n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Genomics↗

Nanopore DNA Sequencing and Genome Assembly on the International Space Station

The emergence of nanopore-based sequencers greatly expands the reach of sequencing into low-resource field environments, enabling in situ molecular analysis. In this work, we evaluated the performance of the MinION DNA sequencer (Oxford Nanopore Technologies) in-flight on the International Space Station (ISS), and benchmarked its performance off-Earth against the MinION, Illumina MiSeq, and PacBio RS II sequencing platforms in terrestrial laboratories. Samples contained mixtures of genomic DNA extracted from lambda bacteriophage, Escherichia coli (strain K12) and Mus musculus (BALB/c). The in-flight sequencing experiments generated more than 80,000 total reads with mean 2D accuracies of 85 - 90%, mean 1D accuracies of 75 - 80%, and median read lengths of approximately 6,000 bases. We were able to construct directed assemblies of the ~4.7 Mb E. coli genome, ~48.5 kb lambda genome, and a representative M. musculus sequence (the ~16.3 kb mitochondrial genome), at 100%, 100%, and 96.7% pairwise identity, respectively, and de novo assemblies of the lambda and E. coli genomes generated solely from nanopore reads yielded 100% and 99.8% genome coverage, respectively, at 100% and 98.5% pairwise identity. Across all surveyed metrics (base quality, throughput, stays/base, skips/base), no observable decrease in MinION performance was observed while sequencing DNA in space. Simulated runs of in-flight nanopore data using an automated bioinformatic pipeline and cloud or laptop based genomic assembly demonstrated the feasibility of real-time sequencing analysis and direct microbial identification in space. Applications of sequencing for space exploration include infectious disease diagnosis, environmental monitoring, evaluating biological responses to spaceflight, and even potentially the detection of extraterrestrial life on other planetary bodies.

Genomics↗