bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

The genome sizes of ostracod crustaceans correlate with body size and phylogeny

Within animals a positive correlation between genome size and body size has been detected in several taxa but not in others, such that it remains unknown how pervasive this pattern may be. Here we provide another example of a positive relationship, in a group of crustaceans whose genome sizes have not previously been investigated. We analyze genome size estimates for 46 species across Class Ostracoda, including 29 new estimates made using Feulgen image analysis densitometry and flow cytometry. Genome sizes in this group range ~80-fold, a level of variability that is otherwise not seen in crustaceans with the exception of some malacostracan orders. We find a strong positive correlation between genome size and body size across all species, including after phylogenetic correction. We additionally detect evidence of XX/XO sex determination in all three species of myodocopids where male and female genome sizes were estimated. On average, genome sizes are larger but less variable in myodocopids than in podocopids, and marine ostracods have larger genomes than freshwater species, but this appears to be explained by phylogenetic inertia. The relationship between phylogeny, genome size, body size, and habitat is complex in this system, and will benefit from additional data collection across various habitats and ostracod taxa.

evolutionary biology↗

A Site Specific Model And Analysis Of The Neutral Somatic Mutation Rate In Whole-Genome Cancer Data

BackgroundDetailed modelling of the neutral mutational process in cancer cells is crucial for identifying driver mutations and understanding the mutational mechanisms that act during cancer development. The neutral mutational process is very complex: whole-genome analyses have revealed that the mutation rate differs between cancer types, between patients and along the genome depending on the genetic and epigenetic context. Therefore, methods that predict the number of different types of mutations in regions or specific genomic elements must consider local genomic explanatory variables. A major drawback of most methods is the need to average the explanatory variables across the entire region or genomic element. This procedure is particularly problematic if the explanatory variable varies dramatically in the element under consideration.\n\nResultsTo take into account the fine scale of the explanatory variables, we model the probabilities of different types of mutations for each position in the genome by multinomial logistic regression. We analyse 505 cancer genomes from 14 different cancer types and compare the performance in predicting mutation rate for both regional based models and site-specific models. We show that for 1000 randomly selected genomic positions, the site-specific model predicts the mutation rate much better than regional based models. We use a forward selection procedure to identify the most important explanatory variables. The procedure identifies site-specific conservation (phyloP), replication timing, and expression level as the best predictors for the mutation rate. Finally, our model confirms and quantifies certain well-known mutational signatures.\n\nConclusionWe find that our site-specific multinomial regression model outperforms the regional based models. The possibility of including genomic variables on different scales and patient specific variables makes it a versatile framework for studying different mutational mechanisms. Our model can serve as the neutral null model for the mutational process; regions that deviate from the null model are candidates for elements that drive cancer development.

bioinformatics↗

riboSeed: leveraging prokaryotic genomic architecture to assemble across ribosomal regions

The vast majority of bacterial genome sequencing has been performed using Illumina short reads. Because of the inherent difficulty of resolving repeated regions with short reads alone, only {approx}10% of sequencing projects have resulted in a closed genome. The most common repeated regions are those coding for ribosomal operons (rDNAs), which occur in a bacterial genome between 1 and 15 times, and are typically used as sequence markers to classify and identify bacteria. Here, we exploit conservation in the genomic context in which rDNAs occur across taxa to improve assembly of these regions relative to de novo sequencing by using the conserved nature of rDNAs across taxa and the uniqueness of their flanking regions within a genome. We describe a method to construct targeted pseudocontigs generated by iteratively assembling reads that map to a reference genomes rDNAs. These pseudocontigs are then used to more accurately assemble the newly-sequenced chromosome. We show that this method, implemented as riboSeed, correctly bridges across adjacent contigs in bacterial genome assembly and, when used in conjunction with other genome polishing tools, can assist in closure of a genome.

bioinformatics↗

Heterogeneity and Intrinsic Variation in Spatial Genome Organization

The genome is hierarchically organized in 3D space and its architecture is altered in differentiation, development and disease. Some of the general principles that determine global 3D genome organization have been established. However, the extent and nature of cell-to-cell and cell-intrinsic variability in genome architecture are poorly characterized. Here, we systematically probe the heterogeneity in genome organization in human fibroblasts by combining high-resolution Hi-C datasets and high-throughput genome imaging. Optical mapping of several hundred genome interaction pairs at the single cell level demonstrates low steady-state frequencies of colocalization in the population and independent behavior of individual alleles in single nuclei. Association frequencies are determined by genomic distance, higher-order chromatin architecture and chromatin environment. These observations reveal extensive variability and heterogeneity in genome organization at the level of single cells and alleles and they demonstrate the coexistence of a broad spectrum of chromatin and genome conformations in a cell population.

molecular biology↗

Genome size and the extinction of small populations

Although extinction is ubiquitous throughout the history of life, insight into the factors that drive extinction events are often difficult to decipher. Most studies of extinction focus on inferring causal factors from past extinction events, but these studies are constrained by our inability to observe extinction events as they occur. Here, we use digital evolution to avoid these constraints and study \"extinction in action\". We focus on the role of genome size in driving population extinction, as previous work both in comparative genomics and digital evolution has shown a correlation between genome size and extinction. We find that extinctions in small populations are caused by large genome size. This relationship between genome size and extinction is due to two genetic mechanisms that increase a populations lethal mutational burden: large genome size leads to both an increased lethal mutation rate and an increased likelihood of stochastic reproduction errors and non-viability. We further show that this increased lethal mutational burden is directly due to genome expansions, as opposed to subsequent adaptation after genome expansion. These findings suggest that large genome size can enhance the extinction likelihood of small populations and may inform which natural populations are at an increased risk of extinction.

evolutionary biology↗

Precision genome-editing with CRISPR/Cas9 in human induced pluripotent stem cells

Genome engineering in human induced pluripotent stem cells (iPSCs) represent an opportunity to examine the contribution of pathogenic and disease modifying alleles to molecular and cellular phenotypes. However, the practical application of genome-editing approaches in human iPSCs has been challenging. We have developed a precise and efficient genome-editing platform that relies on allele-specific guideRNAs (gRNAs) paired with a robust method for culturing and screening the modified iPSC clones. By applying an allele-specific gRNA design strategy, we have demonstrated greatly improved editing efficiency without the introduction of additional modifications of unknown consequence in the genome. Using this approach, we have modified nine independent iPSC lines at five loci associated with neurodegeneration. This genome-editing platform allows for efficient and precise production of isogenic cell lines for disease modeling. Because the impact of CRISPR/Cas9 on off-target sites remains poorly understood, we went on to perform thorough off-target profiling by comparing the mutational burden in edited iPSC lines using whole genome sequencing. The bioinformatically predicted off-target sites were unmodified in all edited iPSC lines. We also found that the numbers of de novo genetic variants detected in the edited and unedited iPSC lines were similar. Thus, our CRISPR/Cas9 strategy does not specifically increase the mutational burden. Furthermore, our analyses of the de novo genetic variants that occur during iPSC culture and genome-editing indicate an enrichment of de novo variants at sites identified in dbSNP. Taken together, we propose that this enrichment represents regions of the genome more susceptible to mutation. Herein, we present an efficient and precise method for allele-specific genome-editing in iPSC and an analyses pipeline to distinguish off-target events from de novo mutations occurring with culture.

genetics↗

A molecular model of the mitochondrial genome segregation machinery in Trypanosoma brucei

In almost all eukaryotes mitochondria maintain their own genome. Despite the discovery more than 50 years ago still very little is known about how the genome is properly segregated during cell division. The protozoan parasite Trypanosoma brucei contains a single mitochondrion with a singular genome the kinetoplast DNA (kDNA). Electron microscopy studies revealed the tripartite attachment complex (TAC) to physically connect the kDNA to the basal body of the flagellum and to ensure proper segregation of the mitochondrial genome via the basal bodies movement, during cell cycle. Using super-resolution microscopy we precisely localize each of the currently known unique TAC components. We demonstrate that the TAC is assembled in a hierarchical order from the base of the flagellum towards the mitochondrial genome and that the assembly is not dependent on the kDNA itself. Based on biochemical analysis the TAC consists of several non-overlapping subcomplexes suggesting an overall size of the TAC exceeding 2.8 mDa. We furthermore demonstrate that the TAC has an impact on mitochondrial organelle positioning however is not required for proper organelle biogenesis or segregation.\n\nSignificance StatementMitochondrial genome replication and segregation are essential processes in most eukaryotic cells. While replication has been studied in some detail much less is known about the molecular machinery required distribute the replicated genomes. Using super-resolution microscopy in combination with molecular biology and biochemistry we show for the first time in which order the segregation machinery is assembled and that it is assembled de novo rather than in a semi conservative fashion in the single celled parasite Trypanosoma brucei. Furthermore, we demonstrate that the mitochondrial genome itself is not required for assembly to occur. It seems that the physical connection of the mitochondrial genome to cytoskeletal elements is a conserved feature in most eukaryotes, however the molecular components are highly diverse.\n\nAbbreviation

cell biology↗

Adaptation in plant genomes: bigger isn’t better, but it’s probably different

Here we have proposed the functional space hypothesis, positing that mutational target size scales with genome size, impacting the number, source, and genomic location of beneficial mutations that contribute to adaptation. Though motivated by preliminary evidence, mostly from Arabidopsis and maize, more data are needed before any rigorous assessment of the hypothesis can be made. If correct, the functional space hypothesis suggests that we should expect plants with large genomes to exhibit more functional mutations outside of genes, more regulatory variation, and likely less signal of strong selective sweeps reducing diversity. These differences have implications for how we study the evolution and development of plant genomes, from where we should look for signals of adaptation to what patterns we expect adaptation to leave in genetic diversity or gene expression data. While flowering plant genomes vary across more than three orders of magnitude in size, most studies of both functional and evolutionary genomics have focused on species at the extreme small edge of this scale. Our hypothesis predicts that methods and results from these small genomes may not replicate well as we begin to explore large plant genomes. Finally, while we have focused here on evidence from plant genomes, we see no a priori reason why similar arguments might not hold in other taxa as well.

evolutionary biology↗

Near-complete Lokiarchaeota genomes from complex environmental samples using long and short read metagenomic analyses

Asgard archaea is a recently proposed superphylum currently comprised of five recognised phyla: Lokiarchaeota, Thorarchaeota, Odinarchaeota, Heimdallarchaeota and Helarchaeota. Members of this group have been identified based on culture-independent approaches with several metagenome-assembled genomes (MAGs) reconstructed to date. However, most of these genomes consist of several relatively small contigs, and, until recently, no complete Asgard archaea genome is yet available. Large scale phylogenetic analyses suggest that Asgard archaea represent the closest archaeal relatives of eukaryotes. In addition, members of this superphylum encode proteins that were originally thought to be specific to eukaryotes, including components of the trafficking machinery, cytoskeleton and endosomal sorting complexes required for transport (ESCRT). Yet, these findings have been questioned on the basis that the genome sequences that underpin them were assembled from metagenomic data, and could have been subjected to contamination and other assembly artefacts. Even though several lines of evidence indicate that the previously reported findings were not affected by these issues, having access to high-quality and preferentially fully closed Asgard archaea genomes is needed to definitively close this debate. Current long-read sequencing technologies such as Oxford Nanopore allow the generation of long reads in a high-throughput manner making them suitable for their use in metagenomics. Although the use of long reads is still limited in this field, recent analyses have shown that it is feasible to obtain complete or near-complete genomes of abundant members of mock communities and metagenomes of various level of complexity. Here, we show that long read metagenomics can be successfully applied to obtain near-complete genomes of low-abundant members of complex communities from sediment samples. We were able to reconstruct six MAGs from different Lokiarchaeota lineages that show high completeness and low fragmentation, with one of them being a near-complete genome only consisting of three contigs. Our analyses confirm that the eukaryote-like features previously associated with Lokiarchaeota are not the result of contamination or assembly artefacts, and can indeed be found in the newly reconstructed genomes.

microbiology↗

A genomic view of coral-associated Prosthecochloris and a companion sulfate-reducing bacterium

Endolithic microbial symbionts in the coral skeleton may play a pivotal role in maintaining coral health. However, compared to aerobic microorganisms, research on the roles of endolithic anaerobic microorganisms and microbe-microbe interactions in the coral skeleton are still in their infancy. In our previous study, we showed that a group of coral-associated Prosthecochloris (CAP), a genus of anaerobic green sulfur bacteria, was dominant in the skeleton of the coral Isopora palifera. Though CAP is diverse, the 16S rRNA phylogeny presents it as a distinct clade separate from other free-living Prosthecochloris. In this study, we build on previous research and further characterize the genomic and metabolic traits of CAP by recovering two new near-complete CAP genomes--Candidatus Prosthecochloris isoporaea and Candidatus Prosthecochloris sp. N1--from coral Isopora palifera endolithic cultures. Genomic analysis revealed that these two CAP genomes have high genomic similarities compared with other Prosthecochloris and harbor several CAP-unique genes. Interestingly, different CAP species harbor various pigment synthesis and sulfur metabolism genes, indicating that individual CAPs can adapt to a diversity of coral microenvironments. A novel near-complete SRB genome--Candidatus Halodesulfovibrio lyudaonia--was also recovered from the same culture. The fact that CAP and various sulfate-reducing bacteria (SRB) co-exist in coral endolithic cultures and coral skeleton highlights the importance of SRB in the coral endolithic community. Based on functional genomic analysis of Ca. P. sp. N1 and Ca. H. lyudaonia, we also propose a syntrophic relationship between the SRB and CAP in the coral skeleton. ImportanceLittle is known about the ecological roles of endolithic microbes in the coral skeleton; one potential role is as a nutrient source for their coral hosts. Here, we identified a close ecological relationship between CAP and SRB. Recovering novel near-complete CAP and SRB genomes from endolithic cultures in this study enabled us to understand the genomic and metabolic features of anaerobic endolithic bacteria in coral skeletons. These results demonstrate that CAP members with similar functions in carbon, sulfur, and nitrogen metabolisms harbor different light-harvesting components, suggesting that CAP in the skeleton adapts to niches with different light intensities. Our study highlights the potential ecological roles of CAP and SRB in coral skeletons and paves the way for future investigations into how coral endolithic communities will respond to environmental changes.

microbiology↗

metaVaR: introducing metavariant species models for reference-free metagenomic-based population genomics

MotivationThe availability of large metagenomic data offers great opportunities for the population genomic analysis of uncultured organisms, especially for small eukaryotes that represent an important part of the unexplored biosphere while playing a key ecological role. However, the majority of these species lacks reference genome or transcriptome which constitutes a technical barrier for classical population genomic analyses. ResultsWe introduce the metavariant species (MVS) model, a representation of the species only by intra-species nucleotide polymorphism. We designed a method combining reference-free variant calling, multiple density-based clustering and maximum weighted independent set algorithms to cluster intra-species variant into MVS directly from multisample metagenomic raw reads without reference genome or reads assembly. The frequencies of the MVS variants are then used to compute population genomic statistics such as FST in order to estimate genomic differentiation between populations and to identify loci under natural selection. The MVSs construction was tested on simulated and real metagenomic data. MVs showed the required quality for robust population genomics and allowed an accurate estimation of genomic differentiation ({Delta}FST < 0.0001 and < 0.03 on simulated and real data respectively). Loci predicted under natural selection on real data were all found by MVSs. MVSs represent a new paradigm that may simplify and enhance holistic approaches for population genomics and evolution of microorganisms. AvailabilityThe method was implemented in a R package, metaVaR. https://github.com/madoui/MetaVaR Contactamadoui@genoscope.cns.fr

bioinformatics↗

Machine learning-based analysis of genomes suggests associations between Wuhan 2019-nCoV and bat Betacoronaviruses

As of February 20, 2020, the 2019 novel coronavirus (renamed to COVID-19) spread to 30 countries with 2130 deaths and more than 75500 confirmed cases. COVID-19 is being compared to the infamous SARS coronavirus, which resulted, between November 2002 and July 2003, in 8098 confirmed cases worldwide with a 9.6% death rate and 774 deaths. Though COVID-19 has a death rate of 2.8% as of 20 February, the 75752 confirmed cases in a few weeks (December 8, 2019 to February 20, 2020) are alarming, with cases likely being under-reported given the comparatively longer incubation period. Such outbreaks demand elucidation of taxonomic classification and origin of the virus genomic sequence, for strategic planning, containment, and treatment. This paper identifies an intrinsic COVID-19 genomic signature and uses it together with a machine learning-based alignment-free approach for an ultra-fast, scalable, and highly accurate classification of whole COVID-19 genomes. The proposed method combines supervised machine learning with digital signal processing for genome analyses, augmented by a decision tree approach to the machine learning component, and a Spearmans rank correlation coefficient analysis for result validation. These tools are used to analyze a large dataset of over 5000 unique viral genomic sequences, totalling 61.8 million bp. Our results support a hypothesis of a bat origin and classify COVID-19 as Sarbecovirus, within Betacoronavirus. Our method achieves high levels of classification accuracy and discovers the most relevant relationships among over 5,000 viral genomes within a few minutes, ab initio, using raw DNA sequence data alone, and without any specialized biological knowledge, training, gene or genome annotations. This suggests that, for novel viral and pathogen genome sequences, this alignment-free whole-genome machine-learning approach can provide a reliable real-time option for taxonomic classification.

bioinformatics↗

Widespread selection against deleterious mutations in the Drosophila genome

We have developed a computational approach to simultaneous genome-wide inference of key population genetics parameters: selection strengths, mutation rates rescaled by the effective population size and the fraction of viable genotypes, solely from an alignment of genomic sequences sampled from the same population. Our approach is based on a generalization of the Ewens sampling formula, used to compute steady-state probabilities of allelic counts in a neutrally evolving population, to populations subjected to selective constraints. Patterns of polymorphisms observed in alignments of genomic sequences are used as input to Approximate Bayesian Computation, which employs the generalized Ewens sampling formula to infer the distributions of population genetics parameters. After carrying out extensive validation of our approach on synthetic data, we have applied it to the evolution of the Drosophila melanogaster genome, where an alignment of 197 genomic sequences is available for a single ancestral-range population from Zambia, Africa. We have divided the Drosophila genome into 100-bp windows and assumed that sequences in each window can exist in either low- or high-fitness state. Thus, the steady-state population in our model is subject to a constant influx of deleterious mutations, which shape the observed frequencies of allelic counts in each window. Our approach, which focuses on deleterious mutations and accounts for intra-window linkage and epistasis, provides an alternative description of background selection. We find that most of the Drosophila genome evolves under selective constraints imposed by deleterious mutations. These constraints are not confined to known functional regions of the genome such as coding sequences and may reflect global biological processes such as the necessity to maintain chromatin structure. Furthermore, we find that inference of mutation rates in the presence of selection leads to mutation rate estimates that are several-fold higher than neutral estimates widely used in the literature. Our computational pipeline can be used in any organism for which a sample of genomic sequences from the same population is available.

evolutionary biology↗

In and outs of Chuviridae endogenous viral elements: origin of a retrovirus and signature of ancient and ongoing arms race in mosquito genomes

BackgroundEndogenous viral elements (EVEs) are sequences of viral origin integrated into the host genome. EVEs have been characterized in various insect genomes, including mosquitoes. A large EVE content has been found in Aedes aegypti and Aedes albopictus genomes among which a recently described Chuviridae viral family is of particular interest, owing to the abundance of EVEs derived from it, the discrepancy in the endogenized gene regions and the frequent association with retrotransposons from the BEL-Pao superfamily. In order to better understand the endogenization process of chuviruses and the association between chuvirus glycoproteins and BEL-Pao retrotransposons, we performed a comparative genomics and evolutionary analysis of chuvirus-derived EVEs found in 37 mosquito genomes. ResultsWe identified 428 EVEs belonging to the Chuviridae family confirming the wide discrepancy between the number of genomic regions endogenized: 409 glycoproteins, 18 RNA-dependent RNA polymerases and one nucleoprotein region. Most of the glycoproteins (263 out of 409) are associated specifically with retroelements from the Pao family. Focusing only on well assembled Pao retroelement copies, we estimated that 263 out of 379 Pao elements are associated with chuvirus-derived glycoproteins. Seventy-three potentially active Pao copies were found to contain glycoproteins into their LTR boundaries. Thirteen out of these were classified as complete and likely autonomous copies, with a full LTR structure and protein domains. We also found 116 Pao copies with no trace of glycoproteins and 37 solo glycoproteins. All potential autonomous Pao copies, contained highly similar LTRs, suggesting a recent/current activity of these elements in the mosquito genomes. ConclusionEvolutionary analysis revealed that most of the glycoproteins found are likely derived from a single or few glycoprotein endogenization events associated with a recombination event with a Pao ancestral element. A potential fully functional Pao-chuvirus hybrid (named Anakin) emerged and the glycoprotein was further replicated through retrotransposition. However, a number of solo glycoproteins, not associated with Pao elements, can still be found in some mosquito genomes 114 million years later, suggesting that these glycoproteins were likely domesticated by the host genome and may participate in an antiviral defense mechanism against both chuvirus and Anakin retrovirus.

evolutionary biology↗

Genomic Prediction with Genotype by Environment Interaction Analysis for Kernel Zinc Concentration in Tropical Maize Germplasm

Zinc (Zn) deficiency is a major risk factor for human health, affecting about 30% of the worlds population. To study the potential of genomic selection (GS) for maize with increased Zn concentration, an association panel and two doubled haploid (DH) populations were evaluated in three environments. Three genomic prediction models, M (M1: Environment + Line, M2: Environment + Line + Genomic, and M3: Environment + Line + Genomic + Genomic x Environment) incorporating main effects (lines and genomic) and the interaction between genomic and environment (G x E) were assessed to estimate the prediction ability (rMP) for each model. Two distinct cross-validation (CV) schemes simulating two genomic prediction breeding scenarios were used. CV1 predicts the performance of newly developed lines, whereas CV2 predicts the performance of lines tested in sparse multi-location trials. Predictions for Zn in CV1 ranged from -0.01 to 0.56 for DH1, 0.04 to 0.50 for DH2 and -0.001 to 0.47 for the association panel. For CV2, rMP values ranged from 0.67 to 0.71 for DH1, 0.40 to 0.56 for DH2 and 0.64 to 0.72 for the association panel. The genomic prediction model which included G x E had the highest average rMP for both CV1 (0.39 and 0.44) and CV2 (0.71 and 0.51) for the association panel and DH2 population, respectively. These results suggest that GS has potential to accelerate breeding for enhanced kernel Zn concentration by facilitating selection of superior genotypes.

genetics↗

Systematic genome-wide querying of coding and non-coding functional elements in E. coli using CRISPRi

Genome-wide repression screens using CRISPR interference (CRISPRi) have enabled the high-throughput identification of essential genes in bacteria. However, there is a lack of functional studies leveraging CRISPRi to systematically explore targeting of both the coding and non-coding genome in bacteria. Here we perform CRISPRi screens in Escherichia coli MG1655 K-12 targeting ~13,000 genomic features, including nearly all protein-coding genes, non-coding RNAs, promoters, and transcription factor binding sites (TFBSs) using a ~33,000-member sgRNA library, which represents the most compact and comprehensive genome-wide CRISPRi library in E. coli to date. Our data reveal insights into the conditional essentiality of the genome with key refinements to screen design and profiling. First, we demonstrate that strong fitness defects associated with essential cellular processes can be resolved using inducible time-series measurements. We show that knockdowns of different classes of genes exhibit distinct, transient responses that are correlated to gene function with genes involved in translation exhibiting the strongest responses. We also query feature essentiality across several biochemical conditions and show that several genes, sRNAs, and operons exhibit conditional phenotypes not reported by previous high-throughput efforts. Second, we evaluate systematically targeting non-genic features (promoters and TFBSs) in the E. coli genome. We show that promoter-targeting guides can be used to add phenotypic confidence to promoter annotations and verify computationally predicted promoters. In contrast to prior studies, we find that promoter knockdowns exhibit a strong targeting orientation dependency where targeting the non-template strand of the promoter closest to the target gene is more effective in knocking down gene expression than other promoter targeting orientations. Unlike eukaryotic genomes, we note that interpreting the effects of TFBS targeting is particularly challenging due to the small size of such features and their proximity to and overlap with other genomic features. Together, this work reveals novel conditionally essential gene phenotypes, provides a characterized set of sgRNAs for future E. coli CRISPRi screens, and highlights considerations for CRISPRi library design and screening for microbial genome characterization.

microbiology↗

Reprogramming the endogenous type III-A CRISPR-Cas system for genome editing, RNA interference and CRISPRi screening in Mycobacterium tuberculosis

Mycobacterium tuberculosis (M.tb) causes the current leading infectious disease. Examination of the functional genomics of M.tb and development of drugs and vaccines are hampered by the complicated and time-consuming genetic manipulation techniques for M.tb. Here, we reprogrammed M.tb endogenous type III-A CRISPR-Cas10 system for simple and efficient gene editing, RNA interference and screening via simple delivery of a plasmid harboring a mini-CRISPR array, thereby avoiding the introduction of exogenous proteins and minimizing proteotoxicity. We demonstrated that M.tb genes were efficiently and specifically knocked-in/out by this system, which was confirmed by whole-genome sequencing. This system was further employed for single and simultaneous multiple-gene RNA interference. Moreover, we successfully applied this system for genome-wide CRISPR interference screening to identify the in-vitro and intracellular growth-regulating genes. This system can be extensively used to explore the functional genomics of M.tb and facilitate the development of new anti-Mycobacterial drugs and vaccines. SummaryTuberculosis caused by Mycobacterium tuberculosis (M.tb) is the current leading infectious disease affecting more than ten million people annually. To dissect the functional genomics and understand its virulence, persistence, and antibiotics resistance, a powerful genome editing tool and high-throughput screening methods are desperately wanted. Our study developed an efficient and a robust tool for genome editing and RNA interference in M.tb using its endogenous CRISPR cas10 system. Moreover, the system has been successfully applied for genome-wide CRISPR interference screening. This tool could be employed to explore the functional genomics of M.tb and facilitate the development of anti-M.tb drugs and vaccines.

microbiology↗

HumGut: A comprehensive Human Gut prokaryotic genomes collection filtered by metagenome data

BackgroundA major bottleneck in the use of metagenome sequencing for human gut microbiome studies has been the lack of a comprehensive genome collection to be used as a reference database. Several recent efforts have been made to re-construct genomes from human gut metagenome data, resulting in a huge increase in the number of relevant genomes. In this work, we aimed to create a collection of the most prevalent healthy human gut prokaryotic genomes, to be used as a reference database, including both MAGs from the human gut and ordinary RefSeq genomes. ResultsWe screened > 5,700 healthy human gut metagenomes for the containment of > 490,000 publicly available prokaryotic genomes sourced from RefSeq and the recently announced UHGG collection. This resulted in a pool of > 379,000 genomes that were subsequently scored and ranked based on their prevalence in the healthy human metagenomes. The genomes were then clustered at subspecies resolution, and cluster representatives were retained to comprise the HumGut collection. Using the Kraken2 software for classification, we find superior performance in the assignment of metagenomic reads, classifying on average 94.5% of the reads in a metagenome, as opposed to 86% with UHGG and 44% when using standard Kraken2 database. HumGut, half the size of standard Kraken2 database and directly comparable to the UHGG size, outperforms them both. ConclusionsThe HumGut collection contains > 30,000 genomes clustered at subspecies resolution and ranked by human gut prevalence. We demonstrate how metagenomes from IBD-patients map equally well to this collection, indicating this reference is relevant also for studies well outside the metagenome reference set used to obtain HumGut. We believe this is a valuable resource in a field in dire need of method standardization. All data and metadata, as well as helpful code, are available at http://arken.nmbu.no/~larssn/humgut/.

microbiology↗