bioRxiv Science⌕ Search

Biology subjects

Hyde, C.

Publications and source records attributed to Hyde, C..

3 recordsLinked to original sources

PfaSTer: A ML-powered serotype caller for Streptococcus pneumoniae genomes

Streptococcus pneumoniae (pneumococcus) is a leading cause of morbidity and mortality worldwide. Although multi-valent pneumococcal vaccines have curbed the incidence of disease, their introduction has resulted in shifted serotype distributions that must be monitored. Whole genome sequence (WGS) data provides a powerful surveillance tool for tracking isolate serotypes, which can be determined from nucleotide sequence of the capsular polysaccharide biosynthetic operon (cps). Although software exists to predict serotypes from WGS data, their use is constrained by the requirement of high-coverage Next Generation Sequencing (NGS) reads. This can present a challenge in so far as accessibility and data sharing. Here we present PfaSTer, a method to identify 65 prevalent serotypes from individual S. pneumoniae genome sequences rather than primary NGS data. PfaSTer combines dimensionality reduction from k-mer analysis with machine learning, allowing for rapid serotype prediction without the need for coverage-based assessments. We then demonstrate the robustness of this method, returning >97% concordance when compared to biochemical results and other in-silico serotypers. PfaSTer is open source and available at: https://github.com/pfizer-opensource/pfaster.

bioinformatics↗

Meta-Analyzed Atopic Dermatitis Transcriptome (MAADT) defines a strong correlation to disease activity and consistent to therapeutic effect

BackgroundAtopic Dermatitis (AD) is a persistent inflammatory disease of the skin to which a few novel treatment options have recently become available. Multiple published datasets, from RNA sequencing (RNA-seq) and microarray experiments performed on lesional (LS) and non-lesional (NL) skin biopsies collected from AD patients, provide a useful resource to better define an AD gene signature and evaluate therapeutic effects. MethodsWe evaluated 22 datasets using defined selection criteria and leave-one-out analysis and then carried out a meta-analysis (M-A) to combine 4 RNA-seq datasets and 5 microarray datasets to define a disease gene signature for AD skin tissue. We used this gene signature to evaluate its correlation to disease activity in published AD datasets, as well as the treatment effect of some of the existing and experimental therapies. ResultsWe report the AD gene signatures developed separately from the RNA-seq or the microarray datasets, as well as a gene signature from datasets combined across these two technologies; all 3 gene signatures showed a strong correlation to the disease activity score (SCORAD) - microarray: Pearsons{rho} = 0.651, p-value < 0.01, RNA-seq:{rho} = 0.640, p-value < 0.01, combined:{rho} = 0.649, p-value < 0.01. The gene signature improvement (GSI) of two existing effective therapies, Dupilumab and Cyclosporine, as well as that of other experimental treatments, is consistent with their reported cohort level efficacy from the associated clinical trials. ConclusionsThe M-A derived AD gene signature provides an evolution of an important resource to correlate gene expression to disease activity and will be helpful for evaluating potential treatment effects for novel therapies.

bioinformatics↗

ScrepYard: an online resource for disulfide-stabilised tandem repeat peptides

Receptor avidity through multivalency is a highly sought-after property of ligands. While readily available in nature in the form of bivalent antibodies, this property remains challenging to engineer in synthetic molecules. The discovery of several bivalent venom peptides containing two homologous and independently folded domains (in a tandem repeat arrangement) has provided a unique opportunity to better understand the underpinning design of multivalency in multimeric biomolecules, as well as how naturally occurring multivalent ligands can be identified. In previous work we classified these molecules as a larger class termed secreted cysteine-rich repeat-proteins (SCREPs). Here, we present an online resource; ScrepYard, designed to assist researchers in identification of SCREP sequences of interest and to aid in characterizing this emerging class of biomolecules. Analysis of sequences within the ScrepYard reveals that two-domain tandem repeats constitute the most abundant SCREP domain architecture, while the interdomain "linker" regions connecting the ordered domains are found to be abundant in amino acids with short or polar sidechains and contain an unusually high abundance of proline residues. Finally, we demonstrate the utility of ScrepYard as a virtual screening tool for discovery of putatively multivalent peptides, by using it as a resource to identify a previously uncharacterised serine protease inhibitor and confirm its predicated activity using an enzyme assay.

bioinformatics↗