bioRxiv ScienceSearch

Biology subjects

Xue, Z.

Publications and source records attributed to Xue, Z..

5 recordsLinked to original sources

Tigmint: Correcting Assembly Errors Using Linked Reads From Large Molecules

Genome sequencing yields the sequence of many short snippets of DNA (reads) from a genome. Genome assembly attempts to reconstruct the original genome from which these reads were derived. This task is difficult due to gaps and errors in the sequencing data, repetitive sequence in the underlying genome, and heterozygosity, and assembly errors are common. These misassemblies may be identified by comparing the sequencing data to the assembly, and by looking for discrepancies between the two. Once identified, these misassemblies may be corrected, improving the quality of the assembly. Although tools exist to identify and correct misassemblies using Illumina pair-end and mate-pair sequencing, no such tool yet exists that makes use of the long distance information of the large molecules provided by linked reads, such as those offered by the 10x Genomics Chromium platform. We have developed the tool Tigmint for this purpose. To demonstrate the effectiveness of Tigmint, we corrected assemblies of a human genome using short reads assembled with ABySS 2.0 and other assemblers. Tigmint reduced the number of misassemblies identified by QUAST in the ABySS assembly by 216 (27%). While scaffolding with ARCS alone more than doubled the scaffold NGA50 of the assembly from 3 to 8 Mbp, the combination of Tigmint and ARCS improved the scaffold NGA50 of the assembly over five-fold to 16.4 Mbp. This notable improvement in contiguity highlights the utility of assembly correction in refining assemblies. We demonstrate its usefulness in correcting the assemblies of multiple tools, as well as in using Chromium reads to correct and scaffold assemblies of long single-molecule sequencing. The source code of Tigmint is available for download from https://github.com/bcgsc/tigmint, and is distributed under the GNU GPL v3.0 license.

genomics

Impact of DNA sequencing and analysis methods on 16S rRNA gene bacterial community analysis in dairy products

DNA sequencing and analysis methods were compared for 16S rRNA V4 PCR amplicon and gDNA mock communities encompassing nine bacterial species commonly found in milk and dairy products. The communities were examined using Illumina MiSeq and Ion Torrent PGM DNA sequencing methods followed by the QIIME 1 (UCLUST) and Divisive Amplicon Denoising Algorithm 2 (DADA2) data analysis pipelines including taxonomic comparisons to the Greengenes and Ribosomal Database Project (RDP) databases. Examination of the PCR amplicon mock community with these methods resulted in Operation Taxonomy Units (OTUs) and Amplicon Sequence Variants (ASVs) that ranged from a low of 13 to high of 118 and were dependent on the DNA sequencing method and read assembly step. The elevated numbers of OTUs and ASVs included assignments to spurious taxa as well as sequence variants of the nine species included in the mock community. Comparisons between the gDNA and PCR amplicon mock communities showed that combining gDNA from the different strains prior to PCR resulted in up to 8.9-fold greater numbers of spurious OTUs and ASVs. However, the DNA sequencing method and initial data assembly steps conferred the largest effects on predictions of bacterial diversity, independent of the mock community type (PCR amplicon or gDNA; Bray-Curtis R2 = 0.88 and weighted Unifrac, R2 = 0.32). Overall, DNA sequencing performed with the Ion Torrent PGM and analyzed with DADA2 and the Greengenes database resulted in the most accurate predictions of the mock community phylogeny, taxonomy, and diversity.\n\nImportanceValidated methods are urgently needed to improve DNA-sequence based assessments of complex bacterial communities. In this study, we used 16S rRNA PCR amplicon and gDNA mock community standards, consisting of nine, dairy-associated bacterial species, to evaluate the most commonly applied 16S rRNA marker gene DNA sequencing and analysis platforms used in evaluating dairy and other bacterial habitats. Our results show that bacterial metataxonomic assessments are largely dependent on the DNA sequencing platform and read curation method used. DADA2 improved sequence annotation compared with QIIME 1, and when combined with the Ion Torrent PGM DNA sequencing platform and the Greengenes database for taxonomic assignment, the most accurate representation of the dairy mock community standards was reached. This approach will be useful for validating sample collection and DNA extraction methods and ultimately investigating bacterial population dynamics in milk and dairy-associated environments.

microbiology

phospho-ERK is a response biomarker to a combination of sorafenib and MEK inhibition in liver cancer

Treatment of liver cancer remains challenging, due to a paucity of drugs that target critical dependencies. Sorafenib is a multikinase inhibitor that is approved as the standard therapy for advanced hepatocellular carcinoma patients, but it can only provide limited survival benefit for patients. To investigate the cause of this limited therapeutic effect, we performed a CRISPR-Cas9 based synthetic lethality screen to search for kinases whose knockout synergize with sorafenib. We find that suppression of ERK2 sensitizes several liver cancer cell lines to sorafenib. Drugs inhibiting the MEK or ERK kinases reverse unresponsiveness to sorafenib in vitro and in vivo in a subset of liver cancer cell lines characterized by high levels of active phospho-ERK levels through synergistic inhibition of ERK kinase activity. Our data provide a combination strategy for treating liver cancer and suggest that tumors with activation of p-ERK, which is seen in some 30% of liver cancers, are most likely to benefit from such combinatorial treatment.

cancer biology

Pan-cancer analysis reveals complex tumor-specific alternative polyadenylation

Alternative polyadenylation (APA) of 3 untranslated regions (3 UTRs) has been implicated in cancer development. Earlier reports on APA in cancer primarily focused on 3 UTR length modifications, and the conventional wisdom is that tumor cells preferentially express transcripts with shorter 3 UTRs. Here, we analyzed the APA patterns of 114 genes, a select list of oncogenes and tumor suppressors, in 9,939 tumor and 729 normal tissue samples across 33 cancer types using RNA-Seq data from The Cancer Genome Atlas, and we found that the APA regulation machinery is much more complicated than what was previously thought. We report 77 cases (gene-cancer type pairs) of differential 3 UTR cleavage patterns between normal and tumor tissues, involving 33 genes in 13 cancer types. For 15 genes, the tumor-specific cleavage patterns are recurrent across multiple cancer types. While the cleavage patterns in certain genes indicate apparent trends of 3 UTR shortening in tumor samples, over half of the 77 cases imply 3 UTR length change trends in cancer that are more complex than simple shortening or lengthening. This work extends the current understanding of APA regulation in cancer, and demonstrates how large volumes of RNA-seq data generated for characterizing cancer cohorts can be mined to investigate this process.

genomics

T-bet+ CD11c+ B Cells Are Critical For Anti-Chromatin IgG Production In The Development Of Lupus

A hallmark of systemic lupus erythematosus is high titers of circulating autoantibody. A novel CD11c+ B cell subset has been identified that is critical for the development of autoimmunity. However, the role of CD11c+ B cells in the development of lupus is unclear. Chronic graft-versus-host disease (cGVHD) is a lupus-like syndrome with great autoantibody production. In the present study we investigated the role of CD11c+ B cells in the pathogenesis of lupus in the cGVHD model. Here, we found the percentage and absolute number of CD11c+ B cells and titer of sera anti-chromatin IgG and IgG2a antibody were increased in cGVHD mice. CD11c+ plasma cells from cGVHD mice produced large amounts of anti-chromatin IgG2a upon stimulation. Depletion of CD11c+ B cells reduced anti-chromatin IgG and IgG2a production. T-bet expression was further shown to be upregulated in CD11c+ B cells. Knockout of T-bet in B cells alleviated cGVHD. The percentage of T-bet+ CD11c+ B cells was elevated in lupus patients and positively correlated with serum anti-chromatin levels. Our findings suggest T-bet+ CD11c+ B cells contribute to the pathogenesis of lupus and provides potential target for therapeutic intervention.

immunology