bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Prioritization of genes driving congenital phenotypes of patients with de novo genomic structural variants

BackgroundGenomic structural variants (SVs) can affect many genes and regulatory elements. Therefore, the molecular mechanisms driving the phenotypes of patients with multiple congenital abnormalities and/or intellectual disability carrying de novo SVs are frequently unknown.\n\nResultsWe applied a combination of systematic experimental and bioinformatic methods to improve the molecular diagnosis of 39 patients with de novo SVs and an inconclusive diagnosis after regular genetic testing. In seven of these cases (18%) whole genome sequencing analysis detected disease-relevant complexities of the SVs missed in routine microarray-based analyses. We developed a computational tool to predict effects on genes directly affected by SVs and on genes indirectly affected due to changes in chromatin organization and impact on regulatory mechanisms. By combining these functional predictions with extensive phenotype information, candidate driver genes were identified in 16/39 (41%) patients. In eight cases evidence was found for involvement of multiple candidate drivers contributing to different parts of the phenotypes. Subsequently, we applied this computational method to a collection of 382 patients with previously detected and classified de novo SVs and identified candidate driver genes in 210 cases (54%), including 32 cases whose SVs were previously not classified as pathogenic. Pathogenic positional effects were predicted in 25% of the cases with balanced SVs and in 8% of the cases with copy number variants.\n\nConclusionsThese results show that driver gene prioritization based on integrative analysis of WGS data with phenotype association and chromatin organization datasets can improve the molecular diagnosis of patients with de novo SVs.

genomics

PyIOmica: Longitudinal Omics Analysis and Classification

SummaryPyIOmica is an open-source Python package focusing on integrating longitudinal multiple omics datasets, characterizing, and classifying temporal trends. The package includes multiple bioinformatics tools including data normalization, annotation, classification, visualization, and enrichment analysis for gene ontology terms and pathways. Additionally, the package includes an implementation of visibility graphs to visualize time series as networks.\n\nAvailability and implementationPyIOmica is implemented as a Python package (pyiomica), available for download and installation through the Python Package Index (PyPI) (https://pypi.python.org/pypi/pyiomica), and can be deployed using the Python import function following installation. PyIOmica has been tested on Mac OS X, Unix/Linux and Microsoft Windows. The application is distributed under an MIT license. Source code for each release is also available for download on Zenodo (https://doi.org/10.5281/zenodo.3342612).\n\nContactgmias@msu.edu

systems biology

H3K4me3 is neither instructive for, nor informed by, transcription.

H3K4me3 is a near-universal histone modification found predominantly at the 5 region of genes, with a well-documented association with gene activity. H3K4me3 has been ascribed roles as both an instructor of gene expression and also a downstream consequence of expression, yet neither has been convincingly proven on a genome-wide scale. Here we test these relationships using a combination of bioinformatics, modelling and experimental data from budding yeast in which the levels of H3K4me3 have been massively ablated. We find that loss of H3K4me3 has no effect on the levels of nascent transcription or transcript in the population. Moreover, we observe no change in the rates of transcription initiation, elongation, mRNA export or turnover, or in protein levels, or cell-to-cell variation of mRNA. Loss of H3K4me3 also has no effect on the large changes in gene expression patterns that follow galactose induction. Conversely, loss of RNA polymerase from the nucleus has no effect on the pattern of H3K4me3 deposition and little effect on its levels, despite much larger changes to other chromatin features. Furthermore, large genome-wide changes in transcription, both in response to environmental stress and during metabolic cycling, are not accompanied by corresponding changes in H3K4me3. Thus, despite the correlation between H3K4me3 and gene activity, neither appear to be necessary to maintain levels of the other, nor to influence their changes in response to environmental stimuli. When we compare gene classes with very different levels of H3K4me3 but highly similar transcription levels we find that H3K4me3-marked genes are those whose expression is unresponsive to environmental changes, and that their histones are less acetylated and dynamically turned-over. Constitutive genes are generally well-expressed, which may alone explain the correlation between H3K4me3 and gene expression, while the biological role of H3K4me3 may have more to do with this distinction in gene class.

genomics

HVEM blockade initiates tumor cell death by innate immunity and improves anti-tumor response by human T cells in NSG immuno-compromised mice

BackgroundTNFRSF14 (herpes virus entry mediator (HVEM) delivers a negative signal to T cells through the B and T Lymphocyte Attenuator (BTLA) molecule and has been associated with a worse prognosis in numerous malignancies. A formal demonstration that the HVEM/BTLA axis can be targeted for cancer immunotherapy is however still lacking. MethodsWe used immunodeficient NOD.SCID.gc-null mice reconstituted with human PBMC and grafted with human tumor cell lines subcutaneously. Tumor growth was compared using linear and non linear regression statistical modeling. The phenotype of tumor-infiltrating leukocytes was determined by flow cytometry. Statistical testing between groups was performed by a non-parametric t test. Quantification of mRNA in the tumor was performed using NanoString pre-designed panels. Bioinformatics analyses were performed using Metascape, Gene Set Enrichment Analysis and Ingenuity Pathways Analysis with embedded statistical testing. ResultsWe showed that a murine monoclonal antibody to human HVEM significantly impacted the growth of various HVEM-positive cancer cell lines in humanized NSG mice. Using CRISPR/cas9 mediated deletion of HVEM, we showed that HVEM expression by the tumor was necessary and sufficient to observe the therapeutic effect. Tumor cell killing by the mAb was dependent on innate immune cells still present in NSG mice, as indicated by in vivo and in vitro assays. Mechanistically, tumor control by human T cells by the mAb was dependent on CD8 T cells and was associated with an increase in the proliferation and number of tumor-infiltrating leukocytes. Accordingly, the expression of genes belonging to T cell activation pathways, such as JAK/STAT and NFKB were enriched in anti-HVEM-treated mice, whereas genes associated with immuno-suppressive pathways were decreased. Finally, we developed a simple in vivo assay to directly demonstrate that HVEM/BTLA is an immune checkpoint for T-cell mediated tumor control. ConclusionsOur results show that targeting HVEM is a promising strategy for cancer immunotherapy.

immunology

Higher quality de novo genome assemblies from degraded museum specimens: a linked-read approach to museomics

AO_SCPLOWBSTRACTC_SCPLOWHigh-throughput sequencing technologies are a proposed solution for accessing the molecular data in historic specimens. However, degraded DNA combined with the computational demands of short-read assemblies has posed significant laboratory and bioinformatics challenges. Linked-read or synthetic long-read sequencing technologies, such as 10X Genomics, may provide a cost-effective alternative solution to assemble higher quality de novo genomes from degraded specimens. Here, we compare assembly quality (e.g., genome contiguity and completeness, presence of orthogroups) between four published genomes assembled from a single shotgun library and four deer mouse (Peromyscus spp.) genomes assembled using 10X Genomics technology. At a similar price-point, these approaches produce vastly different assemblies, with linked-read assemblies having overall higher quality, measured by larger N50 values and greater gene content. Although not without caveats, our results suggest that linked-read sequencing technologies may represent a viable option to build de novo genomes from historic museum specimens, which may prove particularly valuable for extinct, rare, or difficult to collect taxa.

genomics

Mass Spectrometry-based Plasma Proteomics: Considerations from Sample Collection to Achieving Translatable Data

The proteomic analyses of human blood and blood-derived products (e.g. plasma) offers an attractive avenue to translate research progress from the laboratory into the clinic. However, due to its unique protein composition, performing proteomics assays with plasma is challenging. Plasma proteomics has regained interest due to recent technological advances, but challenges imposed by both complications inherent to studying human biology (e.g. inter-individual variability), analysis of biospecimen (e.g. sample variability), as well as technological limitations remain. As part of the Human Proteome Project (HPP), the Human Plasma Proteome Project (HPPP) brings together key aspects of the plasma proteomics pipeline. Here, we provide considerations and recommendations concerning study design, plasma collection, quality metrics, plasma processing workflows, mass spectrometry (MS) data acquisition, data processing and bioinformatic analysis. With exciting opportunities in studying human health and disease though this plasma proteomics pipeline, a more informed analysis of human plasma will accelerate interest whilst enhancing possibilities for the incorporation of proteomics-scaled assays into clinical practice.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=51 SRC=\"FIGDIR/small/716563v2_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (20K):\norg.highwire.dtl.DTLVardef@1faf719org.highwire.dtl.DTLVardef@174c1b1org.highwire.dtl.DTLVardef@5860baorg.highwire.dtl.DTLVardef@366c21_HPS_FORMAT_FIGEXP M_FIG C_FIG

molecular biology

AutoPVS1: An automatic classification tool for PVS1 interpretation of null variants

Null variants are prevalent within human genome, and their accurate interpretation is critical for clinical management. In 2018, the ClinGen Sequence Variant Interpretation (SVI) Working Group refined the only criterion (PVS1) for pathogenicity in the American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP) guidelines. The refinement may improve interpretation consistency, but it also brings hurdles to biocurators because of the complicated workflows and multiple bioinformatics sources required. To address these issues, we developed an automatic classification tool called AutoPVS1 to streamline PVS1 interpretation. We assessed the performance of AutoPVS1 using 56 variants manually curated by ClinGens SVI Working Group and achieved an interpretation concordance of 95% (53/56). A further analysis of 28,586 putative loss-of-function variants by AutoPVS1 demonstrated that at least 27.6% of them do not reach a very strong strength level, with 17.4% based on variant-specific issues and 10.2% on disease mechanism considerations. Moreover, 40.7% (1,918/4,717) of splicing variants were assigned a decreased PVS1 strength level, significantly higher than frameshift and nonsense variants. Our results reinforce the necessity of considering variant-specific issues and disease mechanisms in variant interpretation, and demonstrate that AutoPVS1 is an accurate, reproducible, and reliable tool which facilitates PVS1 interpretation and will thus be of great importance to curators.

genetics

Machine learning based detection of genetic and drug class variant impact on functionally conserved protein binding dynamics

The application of statistical methods to comparatively framed questions about protein dynamics can potentially enable investigations of biomolecular function beyond the current sequence and structural methods in bioinformatics. However, chaotic behavior in single protein trajectories requires statistical inference be obtained from large ensembles of molecular dynamic (MD) simulations representing the comparative functional states of a given protein. Meaningful interpretation of such a complex form of big data poses serious challenges to users of MD. Here, we announce DROIDS v3.0, a molecular dynamic (MD) method + software package for comparative protein dynamics, incorporating many new features including maxDemon v1.0, a multi-method machine learning application that trains on large ensemble comparisons of concerted protein motions in opposing functional states and deploys learned classifications of these states onto newly generated protein dynamic simulations. Local canonical correlations in learning patterns generated from self-similar MD runs are used to identify regions of functionally conserved protein dynamics. Subsequent impacts of genetic and drug class variants on conserved dynamics can also be analyzed by deploying the classifiers on variant MD runs and quantifying how often these altered protein systems display the opposing functional states. Here, we present several case studies of complex changes in functional protein dynamics caused by temperature, genetic mutation, and binding interaction with nucleic acids and small molecules. We studied the impact of genetic variation on functionally conserved protein dynamics in ubiquitin and TATA binding protein and demonstrate that our learning algorithm can properly identify regions of conserved dynamics. We also report impacts to dynamics that correspond well with predicted disruptive effects of a variety of genetic mutations. In addition, we studied the impact of drug class variation on the ATP binding region of Hsp90, similarly identifying conserved dynamics and impacts that rank accordingly with how closely various Hsp90 inhibitors mimic natural ATP binding.\n\nStatement of significanceWe propose a statistical method as well as offer a user-friendly graphical interfaced software pipeline for comparing simulations of the complex motions (i.e. dynamics) of proteins in different functional states. We also provide both method and software to apply artificial intelligence (i.e. machine learning methods) that enable the computer to recognize complex functional differences in protein dynamics on new simulations and report them to the user. This method can identify dynamics important for protein function, as well as to quantify how the motions of molecular variants differ from these important functional dynamic states. For the first time, this method of analysis allows the impacts of different genetic backgrounds or drug classes to be examined within the context of functional motions of the specific protein system under investigation.

biophysics

Pedigree-based measurement of the de novo mutation rate in the gray mouse lemur reveals a high mutation rate, few mutations in CpG sites, and a weak sex bias

Spontaneous germline mutations are the raw material on which evolution acts, and knowledge of their frequency and genomic distribution is crucial for understanding how evolution operates at both long and short timescales. At present, the rate and spectrum of de novo mutations have been directly characterized in only a few lineages. It is therefore critical to expand the phylogenetic scope of these studies to gain a more general understanding of observed mutation rate patterns. Our study provides the first direct mutation rate estimate for a strepsirrhine (i.e., the lemurs and lorises), which comprise nearly half of the primate clade. Using high-coverage linked-read sequencing for a focal quartet of gray mouse lemurs (Microcebus murinus), we estimated the mutation rate to be 1.64 x 10-8 (95% credible interval: 1.41 x 10-8 to 1.98 x 10-8) mutations/site/generation. This estimate is higher than those measured for most previously characterized mammals. Further, we found an unexpectedly low count of paternal mutations, and only a modest overrepresentation of mutations at CpG-sites. Given the surprising nature of these observations, we conducted an independent analysis of context-dependent substitution types for gray mouse lemur and five additional primate species. This analysis yielded patterns consistent with the mutation spectrum from the pedigree mutation-rate analysis, which provides confidence in our ability to accurately identify de novo mutations with our data and bioinformatic filters.

evolutionary biology

Evolutionary dynamics of microbial communities in bioelectrochemical systems.

Bio-electrochemical systems can generate electricity by virtue of mature microbial consortia that gradually and spontaneously optimize performance. To evaluate selective enrichment of these electrogenic microbial communities, five, 3-electrode reactors were inoculated with microbes derived from rice wash wastewater and incubated under a range of applied potentials. Reactors were sampled over a 12-week period and DNA extracted from anodal, cathodal, and planktonic bacterial communities was interrogated using a custom-made bioinformatics pipeline that combined 16S and metagenomic samples to monitor temporal changes in community composition. Some genera that constituted a minor proportion of the initial inoculum dominated within weeks following inoculation and correlated with applied potential. For instance, the abundance of Geobacter increased from 423-fold to 766-fold between -350 mV and -50 mV, respectively. Full metagenomic profiles of bacterial communities were obtained from reactors operating for 12 weeks. Functional analyses of metagenomes revealed metabolic changes between different species of the dominant genus, Geobacter, suggesting that optimal nutrient utilization at the lowest electrode potential is achieved via genome rearrangements and a strong inter-strain selection, as well as adjustment of the characteristic syntrophic relationships. These results reveal a certain degree of metabolic plasticity of electrochemically active bacteria and their communities in adaptation to adverse anodic and cathodic environments.

systems biology

Alcohol Drinking Exacerbates Neural and Behavioral Pathology in the 3xTg-AD Mouse Model of Alzheimer’s Disease

Alzheimers disease (AD) is a progressive neurodegenerative disorder that represents the most common cause of dementia in the United States. Although the link between alcohol use and AD has been studied, preclinical research has potential to elucidate neurobiological mechanisms that underlie this interaction. This study was designed to test the hypothesis that non-dependent alcohol drinking exacerbates the onset and magnitude of AD-like neural and behavioral pathology. We first evaluated the impact of voluntary 24-h, 2-bottle choice home-cage alcohol drinking on the prefrontal cortex and amygdala neuroproteome in C57BL/6J mice and found a striking association between alcohol drinking and AD-like pathology. Bioinformatics identified the AD-associated proteins MAPT (Tau), amyloid beta precursor protein (APP), and presenilin-1 (PSEN-1) as the main modulators of alcohol-sensitive protein networks that included AD-related proteins that regulate energy metabolism (ATP5D, HK1, AK1, PGAM1, CKB), cytoskeletal development (BASP1, CAP1, DPYSL2 [CRMP2], ALDOA, TUBA1A, CFL2, ACTG1), cellular/oxidative stress (HSPA5, HSPA8, ENO1, ENO2), and DNA regulation (PURA, YWHAZ). To address the impact of alcohol drinking on AD, studies were conducted using 3xTg-AD mice that express human MAPT, APP, and PSEN-1 transgenes and develop AD-like brain and behavioral pathology. 3xTg-AD and wildtype mice consumed alcohol or saccharin for 4 months. Behavioral tests were administered during a 1-month alcohol free period. Alcohol intake induced AD-like behavioral pathologies in 3xTg-AD mice including impaired spatial memory in the Morris Water Maze, diminished sensorimotor gating as measured by prepulse inhibition, and exacerbated conditioned fear. Multiplex immunoassay conducted on brain lysates showed that alcohol drinking upregulated primary markers of AD pathology in 3xTg-AD mice: A{beta} 42/40 ratio in the lateral entorhinal and prefrontal cortex and total Tau expression in the lateral entorhinal cortex and amygdala at 1-month post alcohol exposure. Immunocytochemistry showed that alcohol use upregulated expression of pTau (Ser199/Ser202) in the hippocampus, which is consistent with late stage AD. According to the NIA-AA Research Framework, these results suggest that alcohol use is associated with Alzheimers pathology. Results also showed that alcohol use was associated with a general reduction in Akt/mTOR signaling via several phosphoproteins (IR, IRS1, IGF1R, PTEN, ERK, mTOR, p70S6K, RPS6) in multiple brain regions including hippocampus and entorhinal cortex. Dysregulation of Akt/mTOR phosphoproteins suggests alcohol may target this pathway in AD progression. These results suggest that nondependent alcohol drinking increases the onset and magnitude of AD-like neural and behavioral pathology in 3xTg-AD mice.

neuroscience

Global genomic population structure of Clostridioides difficile

Clostridioides difficile is the primary infectious cause of antibiotic-associated diarrhea. Local transmissions and international outbreaks of this pathogen have been previously elucidated by bacterial whole-genome sequencing, but comparative genomic analyses at the global scale were hampered by the lack of specific bioinformatic tools. Here we introduce EnteroBase, a publicly accessible database (http://enterobase.warwick.ac.uk) that automatically retrieves and assembles C. difficile short-reads from the public domain, and calls alleles for core-genome multilocus sequence typing (cgMLST). We demonstrate that the identification of highly related genomes is 89% consistent between cgMLST and single-nucleotide polymorphisms. EnteroBase currently contains 13,515 quality-controlled genomes which have been assigned to hierarchical sets of single-linkage clusters by cgMLST distances. Hierarchical clustering can be used to identify populations of C. difficile at all epidemiological levels, from recent transmission chains through to pandemic and endemic strains, and is largely compatible with prior ribotyping. Hierarchical clustering thus enables comparisons to earlier surveillance data and will facilitate communication among researchers, clinicians and public-health officials who are combatting disease caused by C. difficile.

microbiology

Developmental regulation of Canonical and small ORF translation from mRNA

Ribosomal profiling has revealed the translation of thousands of sequences outside of annotated protein-coding genes, including small Open Reading Frames of less than 100 codons, and the translational regulation of many genes. Here we have improved Poly-Ribo-Seq and applied it to Drosophila melanogaster embryos to extend the catalogue of in-vivo translated small ORFs, and to reveal the translational regulation of both small and canonical ORFs from mRNAs across embryogenesis. We obtain highly correlated samples across five embryonic stages, with close to 500 million putative ribosomal footprints mapped to mRNAs, and compared them to existing Ribo-Seq and proteomic data. Our analysis reveals, for the first time in Drosophila, footprints mapping to codons in a phased pattern, the hallmark of productive translation, and we propose a simple binomial probability metric to ascertain translation probability. However, our results also reveal reproducible ribosomal binding apparently not resulting in productive translation. This non-productive ribosomal binding seems to be especially prevalent amongst upstream short ORFs located in the 5 mRNA Leaders, and amongst canonical ORFs during the activation of the zygotic translatome at the maternal to zygotic transition. We suggest that this non-productive ribosomal binding might be due to cis-regulatory ribosomal binding, and to defective ribosomal scanning of ORFs outside periods of productive translation. Finally, we show that the main function of upstream short ORFs is to buffer the translation of canonical ORFs, and that in general small ORFs in mRNAs display Poly-Ribo-Seq and bioinformatics markers compatible with an evolutionary transitory state towards full coding function.

genomics

Post-EMT: Cadherin-11 mediates cancer hijacking fibroblasts

Current prevailing knowledge on EMT (epithelial mesenchymal transition) deems epithelial cells acquire the characters of mesenchymal cells to be capable of invading and metastasizing on their own. One of the signature events of EMT is called "cadherin switch", e.g. the epithelial E-cadherin switching to the mesenchymal Cadherin-11. Here, we report the critical events after EMT that cancer cells utilize cadherin-11 to hijack the endogenous cadherin-11 positive fibroblasts. Numerous 3-D cell invasion assays with high-content live cell imaging methods reveal that cadherin-11 positive cancer cells adhere to and migrate back and forth dynamically on the cell bodies of fibroblasts. By adhering to fibroblasts for co-invasion through 3-D matrices, cancer cells acquire higher invasion speed and velocity, as well as significantly elevated invasion persistence, which are exclusive characteristics of fibroblast invasion. Silencing cadherin-11 in cancer cells or in fibroblasts, or in both, significantly decouples such physical co-invasion. Additional bioinformatics studies and PDX (patient derived xenograft) studies link such cadherin-11 mediated cancer hijacking fibroblasts to the clinical cancer progression in human such as triple-negative breast cancer patients. Further animal studies confirm cadherin-11 mediates cancer hijacking fibroblasts in vivo and promotes significant solid tumor progression and distant metastasis. Moreover, overexpression of cadherin-11 strikingly protects 4T1-luc cells from implant rejection against firefly luciferase in immunocompetent mice. Overall, our findings report and characterize the critical post-EMT event of cancer hijacking fibroblasts in cancer progression and suggest cadherin-11 can be a therapeutic target for solid tumors with stroma. Our studies hence provide significant updates on the "EMT" theory that EMT cancer cells can hijack fibroblasts to achieve full mesenchymal behaviors in vivo for efficient homing, growth, metastasis and evasion of immune surveillance. Our studies also reveal that cadherin-11 is the key molecule that helps link cancer cells to stromal fibroblasts in the "Seed & Soil" theory. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=117 SRC="FIGDIR/small/729491v2_ufig1.gif" ALT="Figure 1"> View larger version (40K): org.highwire.dtl.DTLVardef@11317f1org.highwire.dtl.DTLVardef@88fffaorg.highwire.dtl.DTLVardef@5da692org.highwire.dtl.DTLVardef@62f6ed_HPS_FORMAT_FIGEXP M_FIG C_FIG

cancer biology

Rapid Detection of Genetic Engineering, Structural Variation, and Antimicrobial Resistance Markers in Bacterial Biothreat Pathogens by Nanopore Sequencing

Widespread release of Bacillus anthracis (anthrax) or Yersinia pestis (plague) would prompt a public health emergency. During an exposure event, high-quality whole genome sequencing (WGS) can identify genetic engineering, including the introduction of antimicrobial resistance (AMR) genes. Here, we developed rapid WGS laboratory and bioinformatics workflows using a long-read nanopore sequencer (MinION) for Y. pestis (6.5h) and B. anthracis (8.5h) and sequenced strains with different AMR profiles. Both salt-precipitation and silica-membrane extracted DNA were suitable for MinION WGS using both rapid and field library preparation methods. In replicate experiments, nanopore quality metrics were defined for genome assembly and mutation analysis. AMR markers were correctly detected and >99% coverage of chromosomes and plasmids was achieved using 100,000 raw sequencing reads. While chromosomes and large and small plasmids were accurately assembled, including novel multimeric forms of the Y. pestis virulence plasmid, pPCP1, MinION reads were error-prone, particularly in homopolymer regions. MinION sequencing holds promise as a practical, front-line strategy for on-site pathogen characterization to speed the public health response during a biothreat emergency.

microbiology

Integrative Genomic Analysis for the Bioprospection of Regulators and Accessory Enzymes Associated with Cellulose Degradation in a Filamentous Fungus (Trichoderma harzianum)

BackgroundUnveiling fungal genome structure and function reveals the potential biotechnological use of fungi. Trichoderma harzianum is a powerful CAZyme-producing fungus. We studied the genomic regions in T. harzianum IOC3844 containing CAZyme genes, transcription factors and transporters.\n\nResultsWe used bioinformatics tools to mine the T. harzianum genome for potential genomics, transcriptomics, and exoproteomics data and coexpression networks. The DNA was sequenced by PacBio SMRT technology for multi-omics data analysis and integration. In total, 1676 genes were annotated in the genomic regions analyzed; 222 were identified as CAZymes in T. harzianum IOC3844. When comparing transcriptome data under cellulose or glucose conditions, 114 genes were differentially expressed in cellulose, with 51 CAZymes. CLR2, a transcription factor physically and phylogenetically conserved in T. harzianum spp., was differentially expressed under cellulose conditions. The genes induced/repressed under cellulose conditions included those important for plant biomass degradation, including CIP2 of the CE15 family and a copper-dependent LPMO of the AA9 family.\n\nConclusionsOur results provide new insights into the relationship between genomic organization and hydrolytic enzyme expression and regulation in T. harzianum IOC3844. Our results can improve plant biomass degradation, which is fundamental for developing more efficient strains and/or enzymatic cocktails for the production of hydrolytic enzymes.

genomics

The role of a priori-identified addiction and smoking gene sets in smoking behaviors.

IntroductionSmoking is a leading cause of death, and genetic variation contributes to smoking behaviors. Identifying genes and sets of genes that contribute to risk for addiction is necessary to prioritize targets for functional characterization and for personalized medicine.\n\nMethodsWe performed a gene set-based association and heritable enrichment study of two addiction-related gene sets, those on the Smokescreen Genotyping Array and the nicotinic acetylcholine receptors, using the largest available GWAS summary statistics. We assessed smoking initiation, cigarettes per day, smoking cessation, and age of smoking initiation.\n\nResultsIndividual genes within each gene set were significantly associated with smoking behaviors. Both sets of genes were significantly associated with cigarettes per day, smoking initiation, and smoking cessation. Age of initiation was only associated with the Smokescreen gene set. While both sets of genes were enriched for trait heritability, each accounts for only a small proportion of the SNP-based heritability (2-12%).\n\nConclusionsThese two gene sets are associated with smoking behaviors, but collectively account for a limited amount of the genetic and phenotypic variation of these complex traits, consistent with high polygenicity.\n\nImplicationsWe evaluated evidence for association and heritable contribution of expert-curated and bioinformatically identified sets of genes related to smoking. Although they impact smoking behaviors, these specifically targeted genes do not account for much of the heritability in smoking and will be of limited use for predictive purposes. Advanced genome-wide approaches and integration of other omics data will be needed to fully account for the genetic variation in smoking phenotypes.

genetics

Detection of genomic alterations in breast cancer with circulating tumour DNA sequencing

Analysis of circulating cell-free DNA (cfDNA) data has opened new opportunities for characterizing tumour mutational landscapes with many applications in genomic-driven oncology. We developed a customized targeted cfDNA sequencing approach for breast cancer (BC) using unique molecular identifiers (UMIs) for error correction. Our assay, spanning a 284.5 kb target region, is combined with freely-available bioinformatics pipelines that provide ultra-sensitive detection of single nucleotide variants (SNVs), and reliable identification of copy number variations (CNVs) directly from plasma DNA. In a cohort of 35 BC patients, our approach detected actionable driver and clonal SNVs at low (~0.5%) frequency levels in cfDNA that were concordant (83.3%) with sequencing of primary and/or metastatic solid tumour sites. We also detected ERRB2 gene CNVs used for HER2 subtype classification with 80% precision compared to immunohistochemistry. Further, we evaluated fragmentation profiles of cfDNA in BC and observed distinct differences compared to data from healthy individuals. Our results show that the developed assay addresses the majority of tumour associated aberrations directly from plasma DNA, and thus may be used to elucidate genomic alterations in liquid biopsy studies.

genomics