bioRxiv ScienceSearch

Biology subjects

Newhouse, S. J.

Publications and source records attributed to Newhouse, S. J..

4 recordsLinked to original sources

ALSgeneScanner: a pipeline for the analysis and interpretation of DNA NGS data of ALS patients

Amyotrophic lateral sclerosis (ALS, MND) is a neurodegenerative disease of upper and lower motor neurons resulting in death from neuromuscular respiratory failure, typically within two years of first symptoms. Genetic factors are an important cause of ALS, with variants in more than 25 genes having strong evidence, and weaker evidence available for variants in more than 120 genes. With the increasing availability of Next-Generation sequencing data, non-specialists, including health care professionals and patients, are obtaining their genomic information without a corresponding ability to analyse and interpret it. Furthermore, the relevance of novel or existing variants in ALS genes is not always apparent. Here we present ALSgeneScanner, a tool that is easy to install and use, able to provide an automatic, detailed, annotated report, on a list of ALS genes from whole genome sequence data in a few hours and whole exome sequence data in about one hour on a readily available mid-range computer. This will be of value to non-specialists and aid in the interpretation of the relevance of novel and existing variants identified in DNA sequencing data.

bioinformatics

DNAscan: a fast, computationally and memory efficient bioinformatics pipeline for the analysis of DNA next-generation-sequencing data

The generation of DNA Next Generation Sequencing (NGS) data is a commonly applied approach for studying the genetic basis of biological processes, including diseases, and underpins the aspirations of precision medicine. However, there are significant challenges when dealing with NGS data. A huge number of bioinformatics tools exist and it is therefore challenging to design an analysis pipeline; NGS analysis is computationally intensive, requiring expensive infrastructure which can be problematic given that many medical and research centres do not have adequate high performance computing facilities and the use of cloud computing facilities is not always possible due to privacy and ownership issues. We have therefore developed a fast and efficient bioinformatics pipeline that allows for the analysis of DNA sequencing data, while requiring little computational effort and memory usage. We achieved this by exploiting state-of-the-art bioinformatics tools. DNAscan can analyse raw, 40x whole genome NGS data in 8 hours, using as little as 8 threads and 16 Gbs of RAM, while guaranteeing a high performance. DNAscan can look for SNVs, small indels, SVs, repeat expansions and viral genetic material (or any other organism). Its results are annotated using a customisable variety of databases including ClinVar, Exac and dbSNP, and a local deployment of the gene.iobio platform is available for an on-the-fly result visualisation.

bioinformatics

Loss of Trem2 in microglia leads to widespread disruption of cell co-expression networks in mouse brain

Rare heterozygous coding variants in the Triggering Receptor Expressed in Myeloid cells 2 (TREM2) gene, conferring increased risk of developing late-onset Alzheimer's disease, have been identified. We examined the transcriptional consequences of the loss of Trem2 in mouse brain to better understand its role in disease using differential expression and coexpression network analysis of Trem2 knockout and wild-type mice. We generated RNA-Seq data from cortex and hippocampus sampled at 4 and 8 months. Using brain cell type markers and ontology enrichment, we found subnetworks with cell type and/or functional identity. We primarily discovered changes in an endothelial-gene enriched subnetwork at 4 months, including a shift towards a more central role for the Amyloid Precursor Protein (App) gene, coupled with widespread disruption of other cell-type subnetworks, including a subnetwork with neuronal identity. We reveal an unexpected potential role of Trem2 in the homeostasis of endothelial cells that goes beyond its known functions as a microglial receptor and signalling hub, suggesting an underlying link between immune response and vascular disease in dementia.

systems biology

Large-Scale Uniform Analysis of Cancer Whole Genomes in Multiple Computing Environments

The International Cancer Genome Consortium (ICGC)s Pan-Cancer Analysis of Whole Genomes (PCAWG) project aimed to categorize somatic and germline variations in both coding and non-coding regions in over 2,800 cancer patients. To provide this dataset to the research working groups for downstream analysis, the PCAWG Technical Working Group marshalled ~800TB of sequencing data from distributed geographical locations; developed portable software for uniform alignment, variant calling, artifact filtering and variant merging; performed the analysis in a geographically and technologically disparate collection of compute environments; and disseminated high-quality validated consensus variants to the working groups. The PCAWG dataset has been mirrored to multiple repositories and can be located using the ICGC Data Portal. The PCAWG workflows are also available as Docker images through Dockstore enabling researchers to replicate our analysis on their own data.

genomics