bioRxiv ScienceSearch

Biology subjects

Hoffman, J.

Publications and source records attributed to Hoffman, J..

5 recordsLinked to original sources

Identification of new therapeutic targets for osteoarthritis through genome-wide analyses of UK Biobank

Osteoarthritis is the most common musculoskeletal disease and the leading cause of disability globally. Here, we perform the largest genome-wide association study for osteoarthritis to date (77,052 cases and 378,169 controls), analysing 4 phenotypes: knee osteoarthritis, hip osteoarthritis, knee and/or hip osteoarthritis, and any osteoarthritis. We discover 64 signals, 52 of them novel, more than doubling the number of established disease loci. Six signals fine map to a single variant. We identify putative effector genes by integrating eQTL colocalization, fine-mapping, human rare disease, animal model, and osteoarthritis tissue expression data. We find enrichment for genes underlying monogenic forms of bone development diseases, and for the collagen formation and extracellular matrix organisation biological pathways. Ten of the likely effector genes, including TGFB1, FGF18, CTSK and IL11 have therapeutics approved or in clinical trials, with mechanisms of action supportive of evaluation for efficacy in osteoarthritis.

genetics

Identification of Novel Common Breast Cancer Risk Variants in Latinas at the 6q25 Locus

Background: Breast cancer is a partially heritable trait and over 180 common genetic variants have been associated with breast cancer in genome wide association studies (GWAS). We have previously performed breast cancer GWAS in Latinas and identified a strongly protective single nucleotide polymorphism (SNP) at 6q25 with the protective minor allele originating from Indigenous American ancestry. Here we report on additional GWAS and replication in Latinas.\n\nMethods: We performed GWAS in 2385 cases and 7342 controls who were either U.S. Latinas or Mexican women. We replicated 2412 cases and 1620 controls of U.S Latina, Mexican, and Colombian women. In addition, we replicated the top novel variants in study of African American and African women and in one study of Chinese women. In each dataset we used logistic regression models to test the association between SNPs and breast cancer risk and corrected for genetic ancestry using either principal components or genetic ancestry inferred from ancestry informative markers using a model based approach.\n\nResults: We identified 3 SNPs (p=1.9x10-8 - 2.8x10-8) at 6q25 locus not in linkage disequilibrium (LD) with variants previously reported at this locus. These SNPs were in high LD with each other, with the top SNP, rs3778609, associated with breast cancer with an odds ratio (OR) and 95% confidence interval (95% CI) of 0.75 (0.68-0.83). In a replication in women of Latin American origin, we also observed a consistent effect (OR: 0.88; 95% CI: 0.78-0.99; p=0.037). Since the minor allele was common in East Asians and African American but not European ancestry populations, we replicated in a meta-analysis of those populations and also observed a consistent effect (OR 0.94; 95% CI: 0.91 - 0.97; p=0.013).\n\nConclusion: The effect size of this variant is relatively large compared to other common variants associated with breast cancer and adds to evidence about the importance of the 6q25 locus for breast cancer susceptibility. Our finding also highlights the utility of performing additional searches for genetic variants for breast cancer in non-European populations.

genetics

RAD sequencing and a hybrid Antarctic fur seal genome assembly reveal rapidly decaying linkage disequilibrium, global population structure and evidence for inbreeding

Recent advances in high throughput sequencing have transformed the study of wild organisms by facilitating the generation of high quality genome assemblies and dense genetic marker datasets. These resources have the potential to significantly advance our understanding of diverse phenomena at the level of species, populations and individuals, ranging from patterns of synteny through rates of linkage disequilibrium (LD) decay and population structure to individual inbreeding. Consequently, we used PacBio sequencing to refine an existing Antarctic fur seal (Arctocephalus gazella) genome assembly and genotyped 83 individuals from six populations using restriction site associated DNA (RAD) sequencing. The resulting hybrid genome comprised 6,169 scaffolds with an N50 of 6.21 Mb and provided clear evidence for the conservation of large chromosomal segments between the fur seal and dog (Canis lupus familiaris). Focusing on the most extensively sampled population of South Georgia, we found that LD decayed rapidly, reaching the background level of r2 = 0.09 by around 26 kb, consistent with other vertebrates but at odds with the notion that fur seals experienced a strong historical bottleneck. We also found evidence for population structuring, with four main Antarctic island groups being resolved. Finally, appreciable variance in individual inbreeding could be detected, reflecting the strong polygyny and site fidelity of the species. Overall, our study contributes important resources for future genomic studies of fur seals and other pinnipeds while also providing a clear example of how high throughput sequencing can generate diverse biological insights at multiple levels of organisation.

evolutionary biology

Multicenter validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU

ObjectivesWe validate a machine learning-based sepsis prediction algorithm (InSight) for detection and prediction of three sepsis-related gold standards, using only six vital signs. We evaluate robustness to missing data, customization to site-specific data using transfer learning, and generalizability to new settings.\n\nDesignA machine learning algorithm with gradient tree boosting. Features for prediction were created from combinations of only six vital sign measurements and their changes over time.\n\nSettingA mixed-ward retrospective data set from the University of California, San Francisco (UCSF) Medical Center (San Francisco, CA) as the primary source, an intensive care unit data set from the Beth Israel Deaconess Medical Center (Boston, MA) as a transfer learning source, and four additional institutions datasets to evaluate generalizability.\n\nParticipants684,443 total encounters, with 90,353 encounters from June 2011 to March 2016 at UCSF.\n\nInterventionsnone\n\nPrimary and secondary outcome measuresArea under the receiver operating characteristic curve (AUROC) for detection and prediction of sepsis, severe sepsis, and septic shock.\n\nResultsFor detection of sepsis and severe sepsis, InSight achieves an area under the receiver operating characteristic (AUROC) curve of 0.92 (95% CI 0.90 - 0.93) and 0.87 (95% CI 0.86 - 0.88), respectively. Four hours before onset, InSight predicts septic shock with an AUROC of 0.96 (95% CI 0.94 -0.98), and severe sepsis with an AUROC of 0.85 (95% CI 0.79 - 0.91).\n\nConclusionsInSight outperforms existing sepsis scoring systems in identifying and predicting sepsis, severe sepsis, and septic shock. This is the first sepsis screening system to exceed an AUROC of 0.90 using only vital sign inputs. InSight is robust to missing data, can be customized to novel hospital data using a small fraction of site data, and retained strong discrimination across all institutions.\n\nStrengths and limitations of this studyO_LIMachine learning is applied to the detection and prediction of three separate sepsis standards in the emergency department, general ward and intensive care settings.\nC_LIO_LIOnly six commonly measured vital signs are used as input for the algorithm.\nC_LIO_LIThe algorithm is robust to randomly missing data.\nC_LIO_LITransfer learning successfully leverages large dataset information to a target dataset.\nC_LIO_LIRetrospective nature of the study does not predict clinician reaction to information.\nC_LI

bioinformatics

Pediatric Severe Sepsis Prediction Using Machine Learning

Early detection of pediatric severe sepsis is necessary in order to administer effective treatment. In this study, we assessed the efficacy of a machine-learning-based prediction algorithm applied to electronic healthcare record (EHR) data for the prediction of severe sepsis onset. The resulting prediction performance was compared with the Pediatric Logistic Organ Dysfunction score (PELOD-2) and pediatric Systemic Inflammatory Response Syndrome score (SIRS) using cross-validation and pairwise t-tests. EHR data were collected from a retrospective set of de-identified pediatric inpatient and emergency encounters drawn from the University of California San Francisco (UCSF) Medical Center, with encounter dates between June 2011 and March 2016. Patients (n = 11,127) were 2-17 years of age and 103 [0.93%] were labeled severely septic. In four-fold cross-validation evaluations, the machine learning algorithm achieved an AUROC of 0.912 for discrimination between severely septic and control pediatric patients at onset and AUROC of 0.727 four hours before onset. Under the same measure, the prediction algorithm also significantly outperformed PELOD-2 (p < 0.05) and SIRS (p < 0.05) in the prediction of severe sepsis four hours before onset. This machine learning algorithm has the potential to deliver high-performance severe sepsis detection and prediction for pediatric inpatients.

bioinformatics