bioRxiv ScienceSearch

Biology subjects

Jones, A. R.

Publications and source records attributed to Jones, A. R..

8 recordsLinked to original sources

ALSgeneScanner: a pipeline for the analysis and interpretation of DNA NGS data of ALS patients

Amyotrophic lateral sclerosis (ALS, MND) is a neurodegenerative disease of upper and lower motor neurons resulting in death from neuromuscular respiratory failure, typically within two years of first symptoms. Genetic factors are an important cause of ALS, with variants in more than 25 genes having strong evidence, and weaker evidence available for variants in more than 120 genes. With the increasing availability of Next-Generation sequencing data, non-specialists, including health care professionals and patients, are obtaining their genomic information without a corresponding ability to analyse and interpret it. Furthermore, the relevance of novel or existing variants in ALS genes is not always apparent. Here we present ALSgeneScanner, a tool that is easy to install and use, able to provide an automatic, detailed, annotated report, on a list of ALS genes from whole genome sequence data in a few hours and whole exome sequence data in about one hour on a readily available mid-range computer. This will be of value to non-specialists and aid in the interpretation of the relevance of novel and existing variants identified in DNA sequencing data.

bioinformatics

Native Mass Spectrometry Reveals the Conformational Diversity of the UVR8 Photoreceptor

UVR8 is a plant photoreceptor protein that regulates photomorphogenic and protective responses to UV light. The inactive, homodimeric state absorbs UV-B light resulting in dissociation into monomers, which are considered to be the active state and comprise a {beta}-propeller core domain and intrinsically disordered N- and C-terminal tails. The C-terminus is required for functional binding to signalling partner COP1. To date, however, structural studies have only been conducted with the core domain where the terminal tails have been truncated. Here, we report structural investigations of full-length UVR8 using native ion mobility mass spectrometry adapted for photo-activation. We show that, whilst truncated UVR8 photo-converts from a single conformation of dimers to a single monomer conformation, the full-length protein exist in numerous conformational families. The full-length dimer adopts both a compact state and an extended state where the C-terminus is primed for activation. In the monomer the extended C-terminus destabilises the core domain to produce highly extended yet stable conformations, which we propose are the fully active states that bind COP1. Our results reveal the conformational diversity of full-length UVR8. We also demonstrate the potential power of native mass spectrometry to probe functionally important structural dynamics of photoreceptor proteins throughout nature.\n\nTOC Graphic\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=141 SRC=\"FIGDIR/small/371658_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (60K):\norg.highwire.dtl.DTLVardef@177c6f9org.highwire.dtl.DTLVardef@a845c9org.highwire.dtl.DTLVardef@17db14forg.highwire.dtl.DTLVardef@103d4f8_HPS_FORMAT_FIGEXP M_FIG C_FIG

biochemistry

Improvements to the rice genome annotation through large-scale analysis of RNA-Seq and proteomics datasets

Rice (Oryza sativa) is one of the most important worldwide crops. The genome has been available for over 10 years and has undergone several rounds of annotation. We created a comprehensive database of transcripts from 29 public RNA sequencing datasets, officially predicted genes from Ensembl plants, and common contaminants in which to search for protein-level evidence. We re-analysed nine publicly accessible rice proteomics datasets. In total, we identified 420K peptide spectrum matches from 47K peptides and 8,187 protein groups. 4168 peptides were initially classed as putative novel peptides (not matching official genes). Following a strict filtration scheme to rule out other possible explanations, we discovered 1,584 high confidence novel peptides. The novel peptides were clustered into 692 genomic loci where our results suggest annotation improvements. 80% of the novel peptides had an ortholog match in the curated protein sequence set from at least one other plant species. For the peptides clustering in intergenic regions (and thus potentially new genes), 101 loci were identified, for which 43 had a high-confidence hit for a protein domain. Our results can be displayed as tracks on the Ensembl genome or other browsers supporting Track Hubs, to support re-annotation of the rice genome.

plant biology

Chromosomally Encoded mcr-5 in Colistin Non-susceptible Pseudomonas aeruginosa

Whole genome sequencing (WGS) of historical Pseudomonas aeruginosa clinical isolates identified a chromosomal copy of mcr-5 within a Tn3-like transposon in P. aeruginosa MRSN 12280. The isolate was non-susceptible to colistin by broth microdilution and genome analysis revealed no mutations known to confer colistin resistance. To the best of our knowledge this is the first report of mcr in colistin non-susceptible P. aeruginosa.

microbiology

Critical assessment of approaches for molecular docking to elucidate associations of HLA alleles with Adverse Drug Reactions

Adverse drug reactions have been linked with genetic polymorphisms in HLA genes in numerous different studies. HLA proteins have an essential role in the presentation of self and non-self peptides, as part of the adaptive immune response. Amongst the associated drugs-allele combinations, anti-HIV drug Abacavir has been shown to be associated with the HLA-B*57:01 allele, and anti-epilepsy drug Carbamazepine with B*15:02, in both cases likely following the altered peptide repertoire model of interaction. Under this model, the drug binds directly to the antigen presentation region, causing different self peptides to be presented, which trigger an unwanted immune response. There is growing interest in searching for evidence supporting this model for other ADRs using bioinformatics techniques. In this study, in silico docking was used to assess the utility and reliability of well-known docking programs when addressing these challenging HLA-drug situations. Four docking programs: SwissDock, ROSIE, AutoDock Vina and AutoDockFR, were used to investigate if each software could accurately dock the Abacavir back into the crystal structure for the protein arising from the known risk allele, and if they were able to distinguish between the HLA-associated and non-HLA-associated (control) alleles. The impact of using homology models on the docking performance and how using different parameters such as including receptor flexibility affected the docking performance, were also investigated to simulate the approach where a crystal structure for a given HLA allele may be unavailable. The programs that were best able to predict the binding position of Abacavir were then used to recreate the docking seen for Carbamazepine with B*15:02 and controls alleles. It was found that the programmes investigated were sometimes able to correctly predict the binding mode of Abacavir with B*57:01 but not always. Each of the software packages that were assessed could predict the binding of Abacavir and Carbamazepine within the correct sub-pocket and, with the exception of ROSIE, was able to correctly distinguish between risk and control alleles. We found that docking to homology models could produce poorer quality predictions, especially when sequence differences impact the architecture of predicted binding pockets. Caution must therefore be used as inaccurate structures may lead to erroneous docking predictions. Incorporating receptor flexibility was found to negatively affect the docking performance for the examples investigated. Taken together, our findings help characterise the potential but also the limitations of computational prediction of drug-HLA interactions. These docking techniques should therefore always be used with care and alongside other methods of investigation, in order to be able to draw strong conclusions from the given results.

bioinformatics

Shared and distinct transcriptomic cell types across neocortical areas

Neocortex contains a multitude of cell types segregated into layers and functionally distinct regions. To investigate the diversity of cell types across the mouse neocortex, we analyzed 12,714 cells from the primary visual cortex (VISp), and 9,035 cells from the anterior lateral motor cortex (ALM) by deep single-cell RNA-sequencing (scRNA-seq), identifying 116 transcriptomic cell types. These two regions represent distant poles of the neocortex and perform distinct functions. We define 50 inhibitory transcriptomic cell types, all of which are shared across both cortical regions. In contrast, 49 of 52 excitatory transcriptomic types were found in either VISp or ALM, with only three present in both. By combining single cell RNA-seq and retrograde labeling, we demonstrate correspondence between excitatory transcriptomic types and their region-specific long-range target specificity. This study establishes a combined transcriptomic and projectional taxonomy of cortical cell types from functionally distinct regions of the mouse cortex.

neuroscience

Extensive non-canonical phosphorylation in human cells revealed using strong-anion exchange-mediated phosphoproteomics

Protein phosphorylation is a ubiquitous post-translational modification (PTM) that regulates all aspects of life. To date, investigation of human cell signalling has focussed on canonical phosphorylation of serine (Ser), threonine (Thr) and tyrosine (Tyr) residues. However, mounting evidence suggests that phosphorylation of histidine also plays a central role in regulating cell biology. Phosphoproteomics workflows rely on acidic conditions for phosphopeptide enrichment, which are incompatible with the analysis of acid-labile phosphorylation such as histidine. Consequently, the extent of non-canonical phosphorylation is likely to be under-estimated.\n\nWe report an Unbiased Phosphopeptide enrichment strategy based on Strong Anion Exchange (SAX) chromatography (UPAX), which permits enrichment of acid-labile phosphopeptides for characterisation by mass spectrometry. Using this approach, we identify extensive and positional phosphorylation patterns on histidine, arginine, lysine, aspartate and glutamate in human cell extracts, including 310 phosphohistidine and >1000 phospholysine sites of protein modification. Remarkably, the extent of phosphorylation on individual non-canonical residues vastly exceeds that of basal phosphotyrosine. Our study reveals the previously unappreciated diversity of protein phosphorylation in human cells, and opens up avenues for exploring roles of acid-labile phosphorylation in any proteome using mass spectrometry.

biochemistry

The proBAM and proBed standard formats: enabling a seamless integration of genomics and proteomics data.

On behalf of The Human Proteome Organization (HUPO) Proteomics Standards Initiative (PSI), we are here introducing two novel standard data formats, proBAM and proBed, that have been developed to address the current challenges of integrating mass spectrometry based proteomics data with genomics and transcriptomics information in proteogenomics studies. proBAM and proBed are adaptations from the well-defined, widely used file formats SAM/BAM and BED respectively, and both have been extended to meet specific requirements entailed by proteomics data. Therefore, existing popular genomics tools such as SAMtools and Bedtools, and several very popular genome browsers, can be used to manipulate and visualize these formats already out-of-the-box. We also highlight that a number of specific additional software tools, properly supporting the proteomics information available in these formats, are now available providing functionalities such as file generation, file conversion, and data analysis. All the related documentation to the formats, including the detailed file format specifications, and example files are accessible at http://www.psidev.info/probam and http://www.psidev.info/probed.

bioinformatics