bioRxiv ScienceSearch

Biology subjects

Patrignani, A.

Publications and source records attributed to Patrignani, A..

4 recordsLinked to original sources

Pushing the limits of de novo genome assembly for complex prokaryotic genomes harboring very long, near identical repeats

Generating a complete, de novo genome assembly for prokaryotes is often considered a solved problem. However, we here show that Pseudomonas koreensis P19E3 harbors multiple, near identical repeat pairs up to 70 kilobase pairs in length. Beyond long repeats, the P19E3 assembly was further complicated by a shufflon region. Its complex genome could not be de novo assembled with long reads produced by Pacific Biosciences technology, but required very long reads from the Oxford Nanopore Technology. Another important factor for a full genomic resolution was the choice of assembly algorithm.\n\nImportantly, a repeat analysis indicated that very complex bacterial genomes represent a general phenomenon beyond Pseudomonas. Roughly 10% of 9331 complete bacterial and a handful of 293 complete archaeal genomes represented this dark matter for de novo genome assembly of prokaryotes. Several of these dark matter genome assemblies contained repeats far beyond the resolution of the sequencing technology employed and likely contain errors, other genomes were closed employing labor-intense steps like cosmid libraries, primer walking or optical mapping. Using very long sequencing reads in combination with assemblers capable of resolving long, near identical repeats will bring most prokaryotic genomes within reach of fast and complete de novo genome assembly.

genomics

CIDER-Seq: unbiased virus enrichment and single-read, full length genome sequencing

Deep-sequencing of virus isolates using short-read sequencing technologies is problematic since viruses are often present in complexes sharing a high-degree of sequence identity. The full-length genomes of such highly-similar viruses cannot be assembled accurately from short sequencing reads. We present a new method, CIDER-Seq (Circular DNA Enrichment Sequencing) which successfully generates accurate full-length virus genomes from individual sequencing reads with no sequence assembly required. CIDER-Seq operates by combining a PCR-free, circular DNA enrichment protocol with Single Molecule Real Time sequencing and a new sequence deconcatenation algorithm. We apply our technique to produce more than 1,200 full-length, highly accurate geminivirus genomes from RNAi-transgenic and control plants in a field trial in Kenya. Using CIDER-Seq we can demonstrate for the first time that the expression of antiviral doublestranded RNA (dsRNA) in transgenic plants causes a consistent shift in virus populations towards species sharing low homology to the transgene derived dsRNA. Our results show that CIDER-seq is a powerful, cost-effective tool for accurately sequencing circular DNA viruses, with future applications in deep-sequencing other forms of circular DNA such as transposons and plasmids.

genomics

An integrative strategy to identify the entire protein coding potential of prokaryotic genomes by proteogenomics

Accurate annotation of all protein-coding sequences (CDSs) is an essential prerequisite to fully exploit the rapidly growing repertoire of completely sequenced prokaryotic genomes. However, large discrepancies among the number of CDSs annotated by different resources, missed functional short open reading frames (sORFs), and overprediction of spurious ORFs represent serious limitations.\n\nOur strategy towards accurate and complete genome annotation consolidates CDSs from multiple reference annotation resources, ab initio gene prediction algorithms and in silico ORFs in an integrated proteogenomics database (iPtgxDB) that covers the entire protein-coding potential of a prokaryotic genome. By extending the PeptideClassifier concept of unambiguous peptides for prokaryotes, close to 95% of the identifiable peptides imply one distinct protein, largely simplifying downstream analysis. Searching a comprehensive Bartonella henselae proteomics dataset against such an iPtgxDB allowed us to unambiguously identify novel ORFs uniquely predicted by each resource, including lipoproteins, differentially expressed and membrane-localized proteins, novel start sites and wrongly annotated pseudogenes. Most novelties were confirmed by targeted, parallel reaction monitoring mass spectrometry, including unique ORFs and variants identified in a re-sequenced laboratory strain that are not present in its reference genome. We demonstrate the general applicability of our strategy for genomes with varying GC content and distinct taxonomic origin, and release iPtgxDBs for B. henselae, Bradyrhozibium diazoefficiens and Escherichia coli as well as the software to generate such proteogenomics search databases for any prokaryote.

genomics

Cell Cycle Constraints and Environmental Control of Local DNA Hypomethylation in Alpha-Proteobacteria

Heritable DNA methylation imprints are ubiquitous and underlie genetic variability from bacteria to humans. In microbial genomes, DNA methylation has been implicated in gene transcription, DNA replication and repair, nucleoid segregation, transposition and virulence of pathogenic strains. Despite the importance of local (hypo)methylation at specific loci, how and when these patterns are established during the cell cycle remains poorly characterized. Taking advantage of the small genomes and the synchronizability of -proteobacteria, we discovered that conserved determinants of the cell cycle transcriptional circuitry establish specific hypomethylation patterns in the cell cycle model system Caulobacter crescentus. We used genome-wide methyl-N6-adenine (m6A-) analyses by restriction-enzyme-cleavage sequencing (REC-Seq) and single-molecule real-time (SMRT) sequencing to show that MucR, a transcriptional regulator that represses virulence and cell cycle genes in S-phase but no longer in G1-phase, occludes 5-GANTC-3 sequence motifs that are methylated by he DNA adenine methyltransferase CcrM. Constitutive expression of CcrM or heterologous methylases in at least two different -proteobacteria homogenizes m6A patterns even when MucR is present and affects promoter activity. Environmental stress (phosphate limitation) can override and reconfigure local hypomethylation patterns imposed by the cell cycle circuitry that dictate when and where local hypomethylation is instated.\n\nAuthor SummaryDNA methylation is the post-replicative addition of a methyl group to a base by a methyltransferase that recognise a specific sequence, and represents an epigenetic regulatory mechanism in both eukaryotes and prokaryotes. In microbial genomes, DNA methylation has been implicated in gene transcription, DNA replication and repair, nucleoid segregation, transposition and virulence of pathogenic strains. CcrM is a conserved, cell cycle regulated adenine methyltransferase that methylates GANTC sites in -proteobacteria. N6-methyl-adenine (m6A) patterns generated by CcrM can change the affinity of a given DNA-binding protein for its target sequence, and therefore affect gene expression. Here, we combine restriction enzyme cleavage-deep sequencing (REC-Seq) with SMRT sequencing to identify hypomethylated 5-GANTC-3 (GANTCs) in -proteobacterial genomes instated by conserved cell cycle factors. By comparing SMRT and REC-Seq data with chromatin immunoprecipitation-deep sequencing data (ChIP-Seq) we show that a conserved transcriptional regulator, MucR, induces local hypomethylation patterns by occluding GANTCs to the CcrM methylase and we provide evidence that this competition occurs during S-phase, but not in G1-phase cells. Furthermore, we find that environmental signals (such as phosphate depletion) are superimposed to the cell cycle control mechanism and can override the specific hypomethylation pattern imposed by the cell cycle transcriptional circuitry.

microbiology