bioRxiv Science⌕ Search

Biology subjects

Machiela, M. J.

Publications and source records attributed to Machiela, M. J..

5 recordsLinked to original sources

LDscore: a scalable, Python 3-powered web platform for LD score regression analysis

Linkage disequilibrium score regression (LDSC) is an important analytical tool for quantifying heritability and estimating genetic correlations between complex traits. However, the LDSC original implementation relies on an outdated Python 2 framework and deploying the standard command-line tools requires significant setup, data access, and computational expertise, creating a barrier for many researchers. To overcome these limitations, we developed LDscore, a significant technical and accessibility upgraded version of LDSC that allows for rapid analysis of GWAS data. The core advancement is the recoding of the LDSC framework in Python 3, enabling computational optimization and ensuring long-term sustainability. Built on top of this improved foundation, LDscore is implemented as a free, publicly available web application integrated within the popular NCI LDlink framework. LDscore can accelerate scientific research by providing an intuitive graphical interface for heritability estimation, genetic correlation, and LD score calculation, including access to an expanded range of reference populations for online analysis. Notably, our results show that selecting the most appropriate reference population LD panel, even at the subcontinental ancestry group level, is essential for minimizing population stratification bias in heritability estimation. By leveraging cloud computing for superior scalability and eliminating the need for local installation, LDscore adheres to FAIR principles, improving access, traceability, and reproducibility across an expanded set of reference populations, and effectively widens access to researchers worldwide providing support for in-depth genetic analyses. Brief summaryLinkage disequilibrium score regression (LDSC), a widely-used method for quantifying heritability and genetic correlation, is limited by an outdated Python 2 framework and complex command-line deployment. We developed LDscore, a significant technical upgrade built on Python 3 for sustainability and computational optimization. LDscore is a free, cloud-based web application integrated into NCI LDlink. LDscore eliminates installation barriers, offering an intuitive interface for computing heritability estimates, LD scores, and genetic correlation. Crucially, LDscore expands the range of reference populations available in LDSC, which can reduce population-stratification-based bias. Leveraging cloud computing, LDscore accelerates and widens global researcher access to LDSC-based genetic computation. AvailabilityLDscore is freely available within LDlink at https://ldlink.nih.gov/ldscore. Source code for the updated LDSC Python3 framework is available at https://github.com/CBIIT/ldsc under the GNU General Public License v3.0 and the webtool code is at https://github.com/CBIIT/nci-webtools-dceg-linkage (webtool code) under the MIT license.

genomics↗

Functional characterization of the 9q34.13 locus identifies RAPGEF1 as modulating risk for melanoma and nevi via RAS activation

Genome-wide association studies identified a melanoma- and nevus count-associated locus on chromosome band 9q34.13. Fine-mapping and melanocyte expression data collectively suggest two potential causal genes with opposite association with risk: higher levels of Rap guanine nucleotide exchange factor 1 (RAPGEF1) and lower levels of uridine-cytidine kinase 1 (UCK1). Colocalization analyses and conditional TWAS suggest multiple causal cis-regulatory sequence variants in partial linkage disequilibrium (LD) to each other. Melanocyte capture-HiC and CRISPR-inhibition demonstrated regulatory interactions between fine-mapped variants and the RAPGEF1 and UCK1 promoters. Focusing on RAPGEF1, we demonstrate RAPGEF1 expression promotes melanocyte growth and drives malignant transformation of human immortalized melanocytes. Following treatment with human EGF, RAPGEF1 overexpression activated both RAP1 and RAS. Further, we show RAPGEF1 expression is significantly enriched in melanomas lacking strongly activating RAS-MAPK mutations, suggesting that RAPGEF1 may promote oncogenic RAS-MAPK signaling in melanomas. Furthermore, in these tumors, we provide preliminary evidence to support the prognostic relevance of RAPGEF1 expression in patients lacking RAS or BRAF mutations. Together with other recent studies, these data suggest that germline variation influencing RAS activation may play a key role in nevus development and melanoma risk.

genetics↗

Integrative multi-omic analysis identifies key transcription factors and target proteins in renal cell carcinoma and its subtypes

To characterize key transcription factors (TFs) whose differential DNA binding can be altered by genetic variants associated with risk for renal cell carcinoma (RCC), we conducted a series of mixed model-based analyses integrating 449 TF ChIP-seq profiles across 9 kidney-related cell lines and summary statistics from a multi-ancestry genome-wide association study of RCC. We identified 96 unique TFs for which presence of SNPs in a neighborhood of TF ChIP-seq peaks are significantly associated (p-value <1x10-4) with their effect on RCC, including EPAS1, ARNT, PAX8 and PBRM1, previously implicated in RCC pathogenesis. Most TFs overlapped active promoters/enhancers in RCC tumors but remained significant after adjusting for tumor chromatin accessibility. Further, we found the co-occupancy of 220 pairs of RCC-related TFs to be associated with RCC risk (FDR<5%) beyond effects of individual TFs, highlighting synergistic regulation between pairs of TFs. To further investigate distal (trans) regulation of TF-binding disruption at RCC associated loci on the proteome, we used a set-based regression to aggregate the trans-effects of multiple loci overlapping with TF binding sites. Across 2,732 proteins profiled in UKB-PPP, identified 169 trans-associated (p-value<1.6x10-7) proteins, nominating specific targets for each TF. For example, we identified TLR3 and ZP3 to be associated with EPAS1, ARNT, and PBRM1, indicating these proteins are likely affected by RCC-related variants disrupting binding sites of the corresponding TFs. These results characterize the landscape of RCC-related TFs and implicate TF-mediated proteomic mechanisms in RCC pathogenesis, nominating testable targets for laboratory studies.

genetics↗

Diploid genome assembly of human fibroblast cell lines enables clone specific variant calling, improved read mapping and accurate phasing

Human cell lines are fundamental tools in biomedical research and are widely used in disease modeling, drug development, and many other domains. Here, we present chromosome-level, phased diploid genome assemblies of two popular human cell lines: the BJ foreskin fibroblast line and the IMR-90 fetal lung fibroblast line. Our high-quality assemblies, generated using long-read and Hi-C sequencing data, reveal substantial structural variation, including more than 50,000 insertions, deletions, duplications, and inversions compared to the recent T2T-CHM13v2.0 reference. Our assemblies provide detailed maps of genetic variation, enabling more accurate variant calling and the ability to phase reads when using newly generated or historical sequencing data on these cell lines or their derivatives. All assemblies and associated data have been made available as a resource for the research community. We envision that diploid genome assembly will become a cornerstone approach for personalized medicine in the near future.

genomics↗

Shared and distinct genetic etiologies for different types of clonal hematopoiesis

Clonal hematopoiesis (CH) - age-related expansion of mutated hematopoietic clones - can differ in frequency and cellular fitness. Descriptive studies have identified a spectrum of events (coding mutations in driver genes (CHIP), gains/losses and copy-neutral loss of chromosomal segments (mCAs), and loss of sex chromosomes). Co-existence of different CH events raises key questions as to their origin, selection, and impact. Here, we report analyses of sequence and genotype array data in up to 482,378 individuals from UK Biobank, demonstrating shared genetic architecture across different types of CH. These data highlighted evidence for a cellular evolutionary trade-off between different forms of CH, with LOY occurring at lower rates in individuals carrying mutations in established CHIP genes. Furthermore, we observed co-occurrence of CHIP and mCAs with overlap at TET2, DNMT3A, and JAK2, in which CHIP precedes mCA acquisition. Individuals carrying these overlapping somatic mutations had a large increase in risk of future hematological malignancy (HR=17.31, 95% CI=9.80-30.58, P=8.94x10-23), which is significantly elevated compared to individuals with non-overlapping CHIP and autosomal mCAs (Pheterogeneity=8.83x10-3). Finally, we leverage the shared genetic architecture of these CH traits to identify 15 novel loci associated with blood cancer risk.

genetics↗