bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Biochemistry”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Using Polysome Isolation with Mechanism Alteration to Uncover Transcriptional and Translational Dynamics in Key Genes

What does it mean when we say a cells biochemistry is regulated during changes to the phenotype? While there are a plethora of potential mechanisms and contributions to the final outcome, a more tractable approach is to examine the dynamics of mRNA. This way, we can assess the contributions of both known and unknown decay and aggregation processes for maintaining levels of gene product on a gene-by-gene basis. In this extended protocol, drug treatments that target specific cellular functions (termed mechanism disruption) can be used in tandem with mRNA extraction from the polysome to look at the dynamics of mRNA levels associated with transcription and translation at multiple stages during a physiological perturbation. This is accomplished through validating the polysome isolation method in human cells and comparing fractions of mRNA for each experimental treatment at multiple points in time. First, three different drug treatments corresponding to the arrest of various cellular processes are administered to populations of human cells. For each treatment, the transcriptome and translatome are compared directly at different time points by assaying both cell-type specific and non-specific genes. There are two findings of note. First, extraction of mRNA from the polysome and comparison with the transcriptome can yield interesting information about the regulation of cellular mRNA during a functional challenge to the cell. In addition, the conventional application of such drugs to assess mRNA decay is an incomplete picture of how severely challenged or senescent cells regulate mRNA in response. This extended protocol demonstrates how the gene- and process-variable degradation of mRNA might ultimately require investigations into the course-grained dynamics of cellular mRNA, from transcription to ribosome.

Cell Biology

Surprisingly weak coordination between leaf structure and function among closely-related tomato species

Natural selection may often favor coordination between different traits, or phenotypic integration, in order to most efficiently acquire and deploy scarce resources. As leaves are the primary photosynthetic organ in plants, many have proposed that leaf physiology, biochemistry, and anatomical structure are coordinated along a functional trait spectrum from fast, resource-acquisitive syndromes to slow, resource-conservative syndromes. However, the coordination hypothesis has rarely been tested at a phylogenetic scale most relevant for understanding rapid adaptation in the recent past or predicting evolutionary trajectories in response to climate change. To that end, we used a common garden to examine genetically-based coordination between leaf traits across 19 wild and cultivated tomato taxa. We found surprisingly weak integration between photosynthetic rate, leaf structure, biochemical capacity, and CO2 diffusion, even though all were arrayed in the predicted direction along a fast-slow spectrum. This suggests considerable scope for unique trait combinations to evolve in response to new environments or in crop breeding. In particular, we find that partially independent variation in stomatal and mesophyll conductance may allow a plant to improve water-use efficiency without necessarily sacrificing maximum photosynthetic rates. Our study does not imply that functional trait spectra or tradeoffs are unimportant, but that the many important axes of variation within a taxonomic group may be unique and not generalizable to other taxa.

Evolutionary Biology

Wikidata as a semantic framework for the Gene Wiki initiative

Open biological data is distributed over many resources making it challenging to integrate, to update and to disseminate quickly. Wikidata is a growing, open community database which can serve this purpose and also provides tight integration with Wikipedia.\n\nIn order to improve the state of biological data, facilitate data management and dissemination, we imported all human and mouse genes, and all human and mouse proteins into Wikidata. In total, 59,530 human genes and 73,130 mouse genes have been imported from NCBI and 27,662 human proteins and 16,728 mouse proteins have been imported from the Swissprot subset of UniProt. As Wikidata is open and can be edited by anybody, our corpus of imported data serves as the starting point for integration of further data by scientists, the Wikidata community and citizen scientists alike. The first use case for this data is to populate Wikipedia Gene Wiki infoboxes directly from Wikidata with the data integrated above. This enables immediate updates of the Gene Wiki infoboxes as soon as the data in Wikidata is modified. Although Gene Wiki pages are currently only on the English language version of Wikipedia, the multilingual nature of Wikidata allows for a usage of the data we imported in all 280 different language Wikipedias. Apart from the Gene Wiki infobox use case, a powerful SPARQL endpoint and up to date exporting functionality (e.g. JSON, XML) enable very convenient further use of the data by scientists.\n\nIn summary, we created a fully open and extensible data resource for human and mouse molecular biology and biochemistry data. This resource enriches all the Wikipedias with structured information and serves as a new linking hub for the biological semantic web.

Bioinformatics

A phylogenetically diverse class of blind type 1 opsins

Opsins are photosensitive proteins catalyzing light-dependent processes across the tree of life. For both microbial (type 1) and metazoan (type 2) opsins, photosensing depends upon covalent interaction between a retinal chromophore and a conserved lysine residue. Despite recent discoveries of potential opsin homologs lacking this residue, phylogenetic dispersal and functional significance of these abnormal sequences have not yet been investigated. We report discovery of a large group of putatively non-retinal binding opsins, present in a number of fungal and microbial genomes and comprising nearly 30% of opsins in the Halobacteriacea, a model clade for opsin photobiology. Based on phylogenetic analyses, structural modeling, genomic context and biochemistry, we propose that these abnormal opsin homologs represent a novel family of sensory opsins which may be involved in taxis response to one or more non-light stimuli. This finding challenges current understanding of microbial opsins as a light-specific sensory family, and provides a potential analogy with the highly diverse signaling capabilities of the eukaryotic G-protein coupled receptors (GPCRs), of which metazoan type 2 opsins are a light-specific sub-clade.

Molecular Biology

IPC - Isoelectric Point Calculator

Accurate estimation of the isoelectric point (pI) based on the amino acid sequence can be useful for many biochemistry and proteomics techniques such as 2-D polyacrylamide gel electrophoresis, or capillary isoelectric focusing used in combination with high-throughput mass spectrometry. Here, I present the Isoelectric Point Calculator, a web service for the estimation of pI using different sets of dissociation constant (pKa) values, including two new, computationally optimized pKa sets. According to the presented benchmarks, IPC outperform previous algorithms by at least 14.9% for proteins and 0.9% for peptides (on average, 22.1% and 59.6%, respectively), which corresponds to an average error of the pI estimation equal to 0.87 and 0.25 pH units for proteins and peptides, respectively. Peptide and protein datasets used in the study and the precalculated pI for PDB, SwissProt databases are available for large-scale analysis and future development. The IPC can be accessed at http://isoelectric.ovh.org.

Bioinformatics

Characterization of a putative NsrR homologue in Streptomyces venezuelae reveals a new member of the Rrf2 superfamily

Members of the Rrf2 superfamily of transcription factors are widespread in bacteria but their functions are largely unexplored. The few that have been characterized in detail sense nitric oxide (NsrR), iron limitation (RirA), cysteine availability (CymR) and the iron sulfur (Fe-S) cluster status of the cell (IscR). In this study we combined ChIP- and dRNA-seq with in vitro biochemistry to characterize a putative NsrR homologue in Streptomyces venezuelae. ChIP-seq analysis revealed that rather than regulating the nitrosative stress response like Streptomyces coelicolor NsrR, Sven6563 binds to a conserved motif at a different, much larger set of genes with a diverse range of functions, including a number of regulators, genes required for glutamine synthesis, NADH/NAD(P)H metabolism, as well as general DNA/RNA and amino acid/protein turn over. Our biochemical experiments further show that Sven6563 has a [2Fe-2S] cluster and that the switch between oxidized and reduced cluster controls its DNA binding activity in vitro. To our knowledge, both the sensing domain and the putative target genes are novel for an Rrf2 protein, suggesting Sven6563 represents a new member of the Rrf2 superfamily. Given the redox sensitivity of its Fe-S cluster we have tentatively named the protein RsrR for Redox sensitive response Regulator.

Molecular Biology

Riparian vegetation limits oxidation of thermodynamically unfavorable bound-carbon stocks along an aquatic interface

In light of increasing terrestrial carbon (C) transport across aquatic boundaries, the mechanisms governing organic carbon (OC) oxidation along terrestrial-aquatic interfaces are crucial to future climate predictions. Here, we investigate the biochemistry, metabolic pathways, and thermodynamics corresponding to OC oxidation in the Columbia River corridor using ultra-high resolution C characterization. We leverage natural vegetative differences to encompass variation in terrestrial C inputs. Our results suggest that decreases in terrestrial C deposition associated with diminished riparian vegetation induce oxidation of physically -bound OC. We also find that contrasting metabolic pathways oxidize OC in the presence and absence of vegetation and--in direct conflict with the priming concept--that inputs of water-soluble and thermodynamically favorable terrestrial OC protects bound-OC from oxidation. In both environments, the most thermodynamically favorable compounds appear to be preferentially oxidized regardless of which OC pool microbiomes metabolize. In turn, we suggest that the extent of riparian vegetation causes sediment microbiomes to locally adapt to oxidize a particular pool of OC, but that common thermodynamic principles govern the oxidation of each pool (i.e., water-soluble or physically-bound). Finally, we propose a mechanistic conceptualization of OC oxidation along terrestrial-aquatic interfaces that can be used to model heterogeneous patterns of OC loss under changing land cover distributions.\n\nKey pointsO_LIRiparian vegetation protects bound-OC stocks\nC_LIO_LIBiochemical processes associated with OC oxidation vary with vegetation conditions\nC_LIO_LICommon thermodynamic principles underlie OC oxidation regardless of vegetation conditions\nC_LI

ecology

Inhibition of DNA2 nuclease as a therapeutic strategy targeting replication stress in cancer cells.

Replication stress is a characteristic feature of cancer cells, which is resulted from sustained proliferative signaling induced by activation of oncogenes or loss of tumor suppressors. In cancer cells, oncogene-induced replication stress manifests as replication-associated lesions, predominantly double-strand DNA breaks (DSBs). An essential mechanism utilized by cells to repair replication-associated DSBs is homologous recombination (HR). In order to overcome replication stress and survive, cancer cells often require enhanced HR repair capacity. Therefore, the key link between HR repair and cellular tolerance to replication-associated DSBs provides us with a mechanistic rationale for exploiting synthetic lethality between HR repair inhibition and replication stress. Our studies showed that DNA2 nuclease is an evolutionarily conserved essential component of HR repair machinery. Here we demonstrate that DNA2 is indeed overexpressed in pancreatic cancers, one of the deadliest and more aggressive forms of human cancers, where mutations in the KRAS are present in 90%-95% of cases. In addition, depletion of DNA2 significantly reduces pancreatic cancer cell survival and xenograft tumor growth, suggesting the therapeutic potential of DNA2 inhibition. Finally, we develop a robust high-throughput biochemistry assay to screen for inhibitors of the DNA2 nuclease activity. The top inhibitors were shown to be efficacious against both yeast Dna2 and human DNA2. Treatment of cancer cells with DNA2 inhibitors recapitulates phenotypes observed upon DNA2 depletion, including decreased DNA end resection and attenuation of HR repair. Similar to genetic ablation of DNA2, chemical inhibition of DNA2 selectively attenuates the growth of various cancer cells with oncogene-induced replication stress. Taken together, our findings open a new avenue to develop a new class of anti-cancer drugs by targeting druggable nuclease DNA2. We 4, 16. In propose DNA2 inhibition as new strategy in cancer therapy by targeting replication stress, a molecular property of cancer cells that is acquired as a result of oncogene activation instead of targeting undruggable oncoprotein itself such as KRAS.

molecular biology

Key Issues Review: Evolution on rugged adaptive landscapes

Adaptive landscapes represent a mapping between genotype and fitness. Rugged adaptive landscapes contain two or more adaptive peaks: allele combinations that differ in two or more genes and confer higher fitness than intermediate combinations. How would a population evolve on such rugged landscapes? Evolutionary biologists have struggled with this question since it was first introduced in the 1930s by Sewall Wright.\n\nDiscoveries in the fields of genetics and biochemistry inspired various mathematical models of adaptive landscapes. The development of landscape models led to numerous theoretical studies analyzing evolution on rugged landscapes under different biological conditions. The large body of theoretical work suggests that adaptive landscapes are major determinants of the progress and outcome of evolutionary processes.\n\nRecent technological advances in molecular biology and microbiology allow experimenters to measure adaptive values of large sets of allele combinations and construct empirical adaptive landscapes for the first time. Such empirical landscapes have already been generated in bacteria, yeast, viruses, and fungi, and are contributing to new insights about evolution on adaptive landscapes.\n\nIn this Key Issues Review we will: (i) introduce the concept of adaptive landscapes; (ii) review the major theoretical studies of evolution on rugged landscapes; (iii) review some of the recently obtained empirical adaptive landscapes; (iv) discuss recent mathematical and statistical analyses motivated by empirical adaptive landscapes, as well as provide the reader with source code and instructions to implement simulations of adaptive landscapes; and (v) discuss possible future directions for this exciting field.

evolutionary biology

An extensible ontology for inference of emergent whole cell function from relationships between subcellular processes

Whole cell responses arise from coordinated interactions between diverse human gene products functioning within various pathways underlying sub-cellular processes (SCP). Lower level SCPs interact to form higher level SCPs, often in a context specific manner to give rise to whole cell function. We sought to determine if capturing such relationships enables us to describe the emergence of whole cell functions from interacting SCPs. We developed the \"Molecular Biology of the Cell\" ontology based on standard cell biology and biochemistry textbooks and review articles. Currently, our ontology contains 5,392 genes, 753 SCPs and 19,182 expertly curated gene-SCP associations. Our algorithm to populate the SCPs with genes enables extension of the ontology on demand and the adaption of the ontology to the continuously growing cell biological knowledge. Since whole cell responses most often arise from the coordinated activity of multiple SCPs, we developed a dynamic enrichment algorithm that flexibly predicts SCP-SCP relationships beyond the current taxonomy. This algorithm enables us to identify interactions between SCPs as a basis for higher order function in a context dependent manner, allowing us to provide a detailed description of how SCPs together can give rise to whole cell functions. We conclude that this ontology can, from omics data sets, enable the development of detailed multidimensional SCP networks for predictive modeling of emergent whole cell functions.

systems biology

Two-Step Interphase Microtubule Disassembly Aids Spindle Morphogenesis

Entry into mitosis triggers profound changes in cell shape and cytoskeletal organisation. Here, by studying microtubule remodelling in human flat mitotic cells, we identify a two-step process of interphase microtubule disassembly. First, a microtubule stabilizing protein, Ensconsin, is inactivated in prophase as a consequence of its phosphorylation downstream of Cdk1/CyclinB. This leads to a reduction in interphase microtubule stability that may help to fuel the growth of centrosomally-nucleated microtubules. The peripheral interphase microtubules that remain are then rapidly lost as the concentration of tubulin heterodimers falls following dissolution of the nuclear compartment boundary. Finally, we show that a failure to destabilize microtubules in prophase leads to the formation of microtubule clumps, which interfere with spindle assembly. Overall, this analysis highlights the importance of the stepwise remodelling of the microtubule cytoskeleton, and the significance of permeabilization of the nuclear envelope in coordinating the changes in cellular organisation and biochemistry that accompany mitotic entry.

cell biology

Auxin Response Factors -- output control in auxin biology

The phytohormone auxin is involved in almost all developmental processes in land plants. Most, if not all, of these processes are mediated by changes in gene expression. Auxin acts on gene expression through a short nuclear pathway that converges upon the activation of a family of DNA-binding transcription factors. These AUXIN RESPONSE FACTORS (ARFs) are thus the effector of auxin response and translate the chemical signal to the regulation of a defined set of genes. Given the limited number of dedicated components in auxin signaling, distinct properties among the ARF family likely contributes to the establishment of multiple unique auxin responses in plant development. In the two decades following the identification of the first ARF in Arabidopsis much has been learnt about how these transcription factors act, and how they generate unique auxin responses. Progress in genetics, biochemistry, genomics and structural biology have helped to develop mechanistic models for ARF action. However, despite intensive efforts, many central questions are yet to be addressed. In this review we highlight what has been learnt about ARF transcription factors, and identify outstanding questions and challenges for the near future.

plant biology

Automated Recommendation Of Metabolite Substructures From Mass Spectra Using Frequent Pattern Mining

Despite the increasing importance of non-targeted metabolomics to answer various life science questions, extracting biochemically relevant information from metabolomics spectral data is still an incompletely solved problem. Most computational tools to identify tandem mass spectra focus on a limited set of molecules of interest. However, such tools are typically constrained by the availability of reference spectra or molecular databases, limiting their applicability to identify unknown metabolites. In contrast, recent advances in the field illustrate the possibility to expose the underlying biochemistry without relying on metabolite identification, in particular via substructure prediction. We describe an automated method for substructure recommendation motivated by association rule mining. Our framework captures potential relationships between spectral features and substructures learned from public spectral libraries. These associations are used to recommend substructures for any unknown mass spectrum. Our method does not require any predefined metabolite candidates, and therefore it can be used for the partial identification of unknown unknowns. The method is called MESSAR (MEtabolite SubStructure Auto-Recommender) and is implemented in a free online web service available at messar.biodatamining.be.\n\nAuthor SummaryMass spectrometry is one of most used techniques to detect and identify metabolites. However, learning metabolite structures directly from mass spectrometry data has always been a challenging task. Thousands of mass spectra from various biological systems still remain unanalyzed simply because no current bioinformatic tools are able to generate structural hypotheses. By manually studying mass spectra of standard compounds, chemists discovered that metabolites that share common substructures can also share spectral features. As data scientists, we believe that such relationships can be unraveled from massive structure and spectra data by machine learning. In this study, we adapted \"association rule mining\", traditionally used in market basket analysis, to structural and spectral data, allowing us to investigate all spectral features - metabolite substructures relationships. We further collected all statistically sound relationships into a database and used them to assign substructral hypotheses to unexplored spectra. We named our approach MESSAR, MEtabolite SubStructure Auto-Recommender, available to the metabolomics and mass spectrometry community as a free and open web service.

bioinformatics

Principles that govern competition or co-existence in Rho-GTPase driven polarization

Rho-GTPases are master regulators of polarity establishment and cell morphology. Positive feedback enables concentration of Rho-GTPases into clusters at the cell cortex, from where they regulate the cytoskeleton. Different cell types reproducibly generate either one (e.g. the front of a migrating cell) or several clusters (e.g. the multiple dendrites of a neuron), but the mechanistic basis for uni-polar or multi-polar outcomes is unclear. The design principles of Rho-GTPase circuits are captured by reaction-diffusion models based on conserved aspects of Rho-GTPase biochemistry. Some such models display rapid winner-takes-all competition between clusters, yielding a unipolar outcome. Other models allow prolonged co-existence of clusters. We derive a \"saturation rule\" general to all relevant models that governs the timescale of competition, and thereby predicts whether the system will generate uni-polar or multi-polar outcomes. We suggest that the saturation rule is a fundamental property of the Rho-GTPase polarity machinery, regardless of the specific feedback mechanism.

cell biology

Simple, single-step, and scar-free mutagenesis of bacterial genes

The need for generating precisely designed mutations is common in genetics, biochemistry, and molecular biology. Here, I describe a new {lambda} Red recombineering method (Direct and Inverted Repeat stimulated excision; DIRex) for fast and easy generation of single point mutations, small insertions or replacements as well as deletions of any size, in bacterial genes. The method does not leave any resistance marker or scar sequence and requires only one transformation to generate a semi-stable intermediate insertion mutant. Spontaneous excision of the intermediate efficiently and accurately generates the final mutant. In addition, the intermediate is transferable between strains by generalized transductions, enabling transfer of the mutation into multiple strains without repeating the recombineering step. Existing methods that can be used to accomplish similar results are either (i) more complicated to design, (ii) more limited in what mutation types can be made, or (iii) require expression of extrinsic factors in addition to {lambda} Red. I demonstrate the utility of the method by generating several deletions, small insertions/replacements, and single nucleotide exchanges in Escherichia coli and Salmonella enterica. Furthermore, the design parameters that influence the excision frequency and the success rate of generating desired point mutations have been examined to determine design guidelines for optimal efficiency.

synthetic biology

Lis1 has two opposing modes of regulating cytoplasmic dynein

Regulation is central to the functional versatility of cytoplasmic dynein, a motor involved in intracellular transport, cell division, and neurodevelopment. Previous work established that Lis1, a conserved and ubiquitous regulator of dynein, binds to its motor domain and induces a tight microtubule-binding state in dynein. The work we present here--a combination of biochemistry, single-molecule assays, cryo-electron microscopy and in vivo experiments--led to the surprising discovery that Lis1 has two opposing modes of regulating dynein, being capable of inducing both low and high affinity for the microtubule. We show that these opposing modes depend on the stoichiometry of Lis1 binding to dynein and that this stoichiometry is regulated by the nucleotide state of dyneins AAA3 domain. We present data on the in vitro and in vivo consequences of abolishing the novel Lis1-induced weak microtubule-binding state in dynein and propose a new model for the regulation of dynein by Lis1.

cell biology

ATP sensing in living plant cells reveals tissue gradients and stress dynamics of energy physiology

Growth and development of plants is ultimately driven by light energy captured through photosynthesis. ATP acts as universal cellular energy cofactor fuelling all life processes, including gene expression, metabolism, and transport. Despite a mechanistic understanding of ATP biochemistry, ATP dynamics in the living plant have been largely elusive. Here we establish live MgATP2- assessment in plants using the fluorescent protein biosensor ATeam1.03-nD/nA. We generate Arabidopsis sensor lines and investigate the sensor in vitro under conditions appropriate for the plant cytosol. We establish an assay for ATP fluxes in isolated mitochondria, and demonstrate that the sensor responds rapidly and reliably to MgATP2- changes in planta. A MgATP2- map of the Arabidopsis seedling highlights different MgATP2- concentrations between tissues and in individual cell types, such as root hairs. Progression of hypoxia reveals substantial plasticity of ATP homeostasis in seedlings, demonstrating that ATP dynamics can be monitored in the living plant.\n\nOne-sentence SummarySensing of MgATP2- by fluorimetry and microscopy allows dissection of ATP fluxes of isolated organelles, and dynamics of cytosolic MgATP2- in vivo.\n\nFunding AgenciesThis work was supported by the Deutsche Forschungsgemeinschaft (DFG) through the Emmy-Noether programme (SCHW1719/1-1; M.S. and GR4251/1-1; C.G.), the Research Training Group GRK 2064 (M.S.; A.J.M.), the Priority Program SPP1710 (A.J.M.) and a grant (SCHW1719/5-1; M.S.) as part of the package PAK918. The Seed Fund grant CoSens from the Bioeconomy Science Center, NRW (A.J.M.; M.S.) is gratefully acknowledged. The scientific activities of the Bioeconomy Science Center were financially supported by the Ministry of Innovation, Science and Research within the framework of the NRW Strategieprojekt BioSC (No. 313/323-400-002 13). A.Co. received funding by the Ministero dellIstruzione, dellUniversita e della Ricerca through the FIRB 2010 programme (RBFR10S1LJ_001) and Piano di Sviluppo di Ateneo 2015 (Universita degli Studi di Milano). M.Z. received funding by the Ministero dellIstruzione, dellUniversita e della Ricerca (Italy) through the PRIN 2010 programme (PRIN2010CSJX4F). S.W. and T.N. received travel support by the Deutscher Akademischer Austauschdienst (DAAD). V.D.C. was supported by the European Social Fund, Operational Programme 2007/2013, and an Erasmus+ Traineeship grant. M.D.F was supported by The Human Frontier Science Program (RPG0053/2012), and the Leverhulme Foundation (RPG-2015-437). I.M.M. was supported by a grant from the Danish Council for Independent Research - Natural Sciences. V.C.P. was supported by the Innovation and Technology Fund (Funding Support to Partner State Key Laboratories in Hong Kong) of the HKSAR.\n\nAbbreviationsAAC - ADP/ATP carrier; AK - adenylate kinase; cAT - carboxyatractyloside; CCCP - carbonyl cyanide m-chlorophenyl hydrazone; CFP - cyan fluorescent protein; CLSM - confocal laser scanning microscopy; ETC - electron transport chain; FRET - Forster Resonance Energy Transfer; LSFM - light sheet fluorescence microscopy.

plant biology

Data Resource Profile: Generation Scotland Electronic Health Record

This paper provides the first detailed demonstration of the research value of the Electronic Health Record (EHR) linked to research data in Generation Scotland Scottish Family Health Study (GS:SFHS) participants, together with how to access this data. The structured, coded variables in the routine biochemistry, prescribing and morbidity records in particular represent highly valuable phenotypic data for a genomics research resource. Access to a wealth of other specialized datasets including cancer, mental health and maternity inpatient information is also possible through the same straightforward and transparent application process. The Electronic Health Record linked dataset is a key component of GS:SFHS, a biobank conceived in 1999 for the purpose of studying the genetics of health areas of current and projected public health importance. Over 24,000 adults were recruited from 2006 to 2011, with broad and enduring written informed consent for biomedical research. Consent was obtained from 23,603 participants for GS:SFHS study data to be linked to their Scottish National Health Service (NHS) records, using their Community Health Index (CHI) number. This identifying number is used for NHS Scotland procedures (registrations, attendances, samples, prescribing and investigations) and allows healthcare records for individuals to be linked across time and location. Here, we describe the NHS EHR dataset on the sub-cohort of 20,032 GS:SFHS participants with consent and mechanism for record linkage plus extensive genetic data. Together with existing study phenotypes, including family history and environmental exposures such as smoking, the EHR is a rich resource of real world data that can be used in research to characterise the health trajectory of participants, available at low cost and a high degree of timeliness, matched to DNA, urine and serum samples and genome-wide genetic information.

genetics