bioRxiv ScienceSearch

Biology subjects

Wright, J.

Publications and source records attributed to Wright, J..

11 recordsLinked to original sources

Cohort Profile: East London Genes & Health (ELGH), a community based population genomics and health study in people of British-Bangladeshi and -Pakistani heritage.

Cohort profile in a nutshellO_LIEast London Genes & Health (ELGH) is a large scale, community genomics and health study (to date >34,000 volunteers; target 100,000 volunteers). C_LIO_LIELGH was set up in 2015 to gain deeper understanding of health and disease, and underlying genetic influences, in British-Bangladeshi and British-Pakistani people living in east London. C_LIO_LIELGH prioritises studies in areas important to, and identified by, the community it represents. Current priorities include cardiometabolic diseases and mental illness, these being of notably high prevalence and severity. However studies in any scientific area are possible, subject to community advisory group and ethical approval. C_LIO_LIELGH combines health data science (using linked UK National Health Service (NHS) electronic health record data) with exome sequencing and SNP array genotyping to elucidate the genetic influence on health and disease, including the contribution from high rates of parental relatedness on rare genetic variation and homozygosity (autozygosity), in two understudied ethnic groups. Linkage to longitudinal health record data enables both retrospective and prospective analyses. C_LIO_LIThrough Stage 2 studies, ELGH offers researchers the opportunity to undertake recall-by-genotype and/or recall-by-phenotype studies on volunteers. Sub-cohort, trial-within-cohort, and other study designs are possible. C_LIO_LIELGH is a fully collaborative, open access resource, open to academic and life sciences industry scientific research partners. C_LI

genomics

An Artificial Intelligence Workflow for Defining Host-Pathogen Interactions

For image-based infection biology, accurate unbiased quantification of host-pathogen interactions is essential, yet often performed manually or using limited enumeration employing simple image analysis algorithms based on image segmentation. Host protein recruitment to pathogens is often refractory to accurate automated assessment due to its heterogeneous nature. An intuitive intelligent image analysis program to assess host protein recruitment within general cellular pathogen defense is lacking. We present HRMAn (Host Response to Microbe Analysis), an open-source image analysis platform based on machine learning algorithms and deep learning. We show that HRMAn has the capability to learn phenotypes from the data, without relying on researcher-based assumptions. Using Toxoplasma gondii and Salmonella typhimurium we demonstrate HRMAns capacity to recognize, classify and quantify pathogen killing, replication and cellular defense responses.

microbiology

Short-term insurance versus long-term bet-hedging strategies as adaptations to variable environments

Understanding how organisms adapt to environmental variation is a key challenge of biology. Central to this are bet-hedging strategies that maximize geometric mean fitness across generations, either by being conservative or diversifying phenotypes. Theoretical models of bet-hedging and the multiplicative fitness effects of environmental variation across generations have traditionally assumed that environmental conditions are constant within lifetimes. However, behavioral ecology has revealed adaptive responses to additive fitness effects of environmental variation within lifetimes, either through insurance or risk-sensitive strategies. Here we explore whether the effects of adaptive insurance interact with the evolution of bet-hedging by varying the position and skew of fitness functions within and between lifetimes. When insurance causes the optimal phenotype to shift from the peak to down the less steeply decreasing side of the fitness function, then conservative bet-hedging does not generally evolve on top of this, even if diversifying bet-hedging can. Canalization to reduce phenotypic variation within a lifetime is almost always favored, except when the tails of the fitness function are steeply convex and produce a novel risk-sensitive increase in phenotypic variance akin to diversifying bet-hedging. Importantly, using skewed fitness functions, we provide the first example of how conservative and diversifying bet-hedging strategies might coexist.

evolutionary biology

Nearly all new protein-coding predictions in the CHESS database are not protein-coding

In a 2018 paper posted to bioRxiv, Pertea et al. presented the CHESS database, a new catalog of human gene annotations that includes 1,178 new protein-coding predictions. These are based on evidence of transcription in human tissues and homology to earlier annotations in human and other mammals. Here, we reanalyze the evidence used by CHESS, and find that nearly all protein-coding predictions are false positives. We find that 86% overlap transposons marked by RepeatMasker that are known to frequently result in false positive protein-coding predictions. More than half are homologous to only nine Alu-derived primate sequences corresponding to an erroneous and previously withdrawn Pfam protein domain. The entire set shows poor evolutionary conservation and PhyloCSF protein-coding evolutionary signatures indistinguishable from noncoding RNAs, indicating lack of protein-coding constraint. Only four predictions are supported by mass spectrometry evidence, and even those matches are inconclusive. Overall, the new protein-coding predictions are unsupported by any credible experimental or evolutionary evidence of function, result primarily from homology to genes incorrectly classified as protein-coding, and are unlikely to encode functional proteins.

genomics

FMRIPrep: a robust preprocessing pipeline for functional MRI

Preprocessing of functional MRI (fMRI) involves numerous steps to clean and standardize data before statistical analysis. Generally, researchers create ad hoc preprocessing workflows for each new dataset, building upon a large inventory of tools available for each step. The complexity of these workflows has snowballed with rapid advances in MR data acquisition and image processing techniques. We introduce fMRIPrep, an analysis-agnostic tool that addresses the challenge of robust and reproducible preprocessing for task-based and resting fMRI data. FMRIPrep automatically adapts a best-in-breed workflow to the idiosyncrasies of virtually any dataset, ensuring high-quality preprocessing with no manual intervention. By introducing visual assessment checkpoints into an iterative integration framework for software-testing, we show that fMRIPrep robustly produces high-quality results on a diverse fMRI data collection comprising participants from 54 different studies in the OpenfMRI repository. We review the distinctive features of fMRIPrep in a qualitative comparison to other preprocessing workflows. We demonstrate that fMRIPrep achieves higher spatial accuracy as it introduces less uncontrolled spatial smoothness than commonly used preprocessing tools. FMRIPrep has the potential to transform fMRI research by equipping neuroscientists with a high-quality, robust, easy-to-use and transparent preprocessing workflow which can help ensure the validity of inference and the interpretability of their results.

bioinformatics

Skip-mers: increasing entropy and sensitivity to detect conserved genic regions with simple cyclic q-grams

Bioinformatic analyses and tools make extensive use of k-mers (fixed contiguous strings of k nucleotides) as an informational unit. K-mer analyses are both useful and fast, but are strongly affected by single-nucleotide polymorphisms or sequencing errors, effectively hindering direct-analyses of whole regions and decreasing their usability between evolutionary distant samples.\n\nWe introduce a concept of skip-mers, a cyclic pattern of used-and-skipped positions of k nucleotides spanning a region of size S [≥] k, and show how analyses are improved compared to using k-mers. The entropy of skip-mers increases with the larger span, capturing information from more distant positions and increasing the specificity, and uniqueness, of larger span skip-mers within a genome. In addition, skip-mers constructed in cycles of 1 or 2 nucleotides in every 3 (or a multiple of 3) lead to increased sensitivity in the coding regions of genes, by grouping together the more conserved nucleotides of the protein-coding regions.\n\nWe implemented a set of tools to count and intersect skip-mers between different datasets. We used these tools to show how skip-mers have advantages over k-mers in terms of entropy and increased sensitivity to detect conserved coding sequence, allowing better identification of genic matches between evolutionarily distant species. We also highlight potential applications to problems such as whole-genome alignment and multi-genome evolutionary analyses.\n\nSoftware availabilitythe skm-tools implementing the methods described in this manuscript are available under MIT license at http://github.com/bioinfologics/skm-tools/

bioinformatics

CRISPR/Cas9-APEX-mediated proximity labeling enables discovery of proteins associated with a predefined genomic locus in living cells

The activation or repression of a genes expression is primarily controlled by changes in the proteins that occupy its regulatory elements. The most common method to identify proteins associated with genomic loci is chromatin immunoprecipitation (ChIP). While having greatly advanced our understanding of gene expression regulation, ChIP requires specific, high quality, IP-competent antibodies against nominated proteins, which can limit its utility and scope for discovery. Thus, a method able to discover and identify proteins associated with a particular genomic locus within the native cellular context would be extremely valuable. Here, we present a novel technology combining recent advances in chemical biology, genome targeting, and quantitative mass spectrometry to develop genomic locus proteomics, a method able to identify proteins which occupy a specific genomic locus.

biochemistry

The ash dieback invasion of Europe was founded by two individuals from a native population with huge adaptive potential

Accelerating international trade and climate change make pathogen spread an increasing concern. Hymenoscyphus fraxineus, the causal agent of ash dieback is one such pathogen, moving across continents and hosts from Asian to European ash. Most European common ash (Fraxinus excelsior) trees are highly susceptible to H. fraxineus although a small minority (~5%) evidently have partial resistance to dieback. We have assembled and annotated a draft of the H. fraxineus genome which approaches chromosome scale. Pathogen genetic diversity across Europe, and in Japan, reveals a tight bottleneck into Europe, though a signal of adaptive diversity remains in key host interaction genes (effectors). We find that the European population was founded by two divergent haploid individuals. Divergence between these haplotypes represents the 'shadow' of a large source population and subsequent introduction would greatly increase adaptive potential and the pathogen's threat. Thus, EU wide biological security measures remain an important part of the strategy to manage this disease.

genomics

Single Cell Transcriptomics And Flow Cytometry Reveal Disease-Associated Fibroblast Subsets In Rheumatoid Arthritis

Fibroblasts mediate normal tissue matrix remodeling, but they can cause fibrosis or tissue destruction following chronic inflammation. In rheumatoid arthritis (RA), synovial fibroblasts expand, degrade cartilage, and drive joint inflammation. Little is known about fibroblast heterogeneity or if aberrations in fibroblast subsets relate to disease pathology. Here, we used an integrative strategy, including bulk transcriptomics on targeted subpopulations and unbiased single-cell transcriptomics, to analyze fibroblasts from synovial tissues. We identify 7 phenotypic fibroblast subsets with distinct surface protein phenotypes, and these collapsed into 3 subsets based on transcriptomics data. One subset expressing PDPN, THY1, but lacking CD34 was 3-fold expanded in RA relative to osteoarthritis (P=0.007); most of these cells expressed CDH11. The subsets were found to differ in expression of cytokines and matrix metalloproteinases, localization in synovial microanatomy, and in response to TNF. Our approach provides a template to identify pathogenic stromal cellular subsets in complex diseases.

immunology

W2RAP: a pipeline for high quality, robust assemblies of large complex genomes from short read data

Producing high-quality whole-genome shotgun de novo assemblies from plant and animal species with large and complex genomes using low-cost short read sequencing technologies remains a challenge. But when the right sequencing data, with appropriate quality control, is assembled using approaches focused on robustness of the process rather than maximization of a single metric such as the usual contiguity estimators, good quality assemblies with informative value for comparative analyses can be produced. Here we present a complete method described from data generation and qc all the way up to scaffold of complex genomes using Illumina short reads and its application to data from plants and human datasets. We show how to use the w2rap pipeline following a metric-guided approach to produce cost-effective assemblies. The assemblies are highly accurate, provide good coverage of the genome and show good short range contiguity. Our pipeline has already enabled the rapid, cost-effective generation of de novo genome assemblies from large, polyploid crop species with a focus on comparative genomics.\n\nAvailabilityw2rap is available under MIT license, with some subcomponents under GPL-licenses. A ready-to-run docker with all software pre-requisites and example data is also available.\n\nhttp://github.com/bioinfologics/w2rap\n\nhttp://github.com/bioinfologics/w2rap-contigger

bioinformatics

An improved assembly and annotation of the allohexaploid wheat genome identifies complete families of agronomic genes and provides genomic evidence for chromosomal translocations.

Advances in genome sequencing and assembly technologies are generating many high quality genome sequences, but assemblies of large, repeat-rich polyploid genomes, such as that of bread wheat, remain fragmented and incomplete. We have generated a new wheat whole-genome shotgun sequence assembly using a combination of optimised data types and an assembly algorithm designed to deal with large and complex genomes. The new assembly represents more than 78% of the genome with a scaffold N50 of 88.8kbp that has a high fidelity to the input data. Our new annotation combines strand-specific Illumina RNAseq and PacBio full-length cDNAs to identify 104,091 high confidence protein-coding genes and 10,156 non-coding RNA genes. We confirmed three known and identified one novel genome rearrangements. Our approach enables the rapid and scalable assembly of wheat genomes, the identification of structural variants, and the definition of complete gene models, all powerful resources for trait analysis and breeding of this key global crop. [Supplemental material is available for this article.]

genomics