bioRxiv Science⌕ Search

Biology subjects

Mews, M.

Publications and source records attributed to Mews, M..

3 recordsLinked to original sources

Evaluating sequence-to-function deep learning models for ancestry-stratified regulatory variant effect prediction using multi-ancestry blood eQTLs

BackgroundSequence-to-function (S2F) deep learning models are increasingly used to prioritize non-coding regulatory variants, but their behavior across ancestrally diverse populations remains unclear. Because both training data and reference resources are heavily European-centered, multi-ancestry benchmarks are needed to determine whether S2F scores capture regulatory effects consistently across populations with different allele-frequency and LD patterns. MethodsWe evaluated Borzoi and AlphaGenome using whole blood eQTL data from the MAGENTA cohort, including African American (AA; N = 224), Caribbean Hispanic (CH; N = 209), and Non-Hispanic White (NHW; N = 235) participants. Model predictions were benchmarked against sampled nominal eQTLs and ancestry-stratified SuSiE fine-mapped variants using Spearman correlation, direction concordance, inter-model convergence, and distance-matched AUROC, with sensitivity analyses for minor allele frequency and comparison-set definition. We also compared FILER functional annotation overlap among high-Posterior Inclusion Probability (PIP) variants across ancestries. ResultsBoth models showed weak agreement with nominal eQTL effect sizes across ancestries and TSS-distance bins ({rho} [&le;] 0.138), with direction concordance only marginally above chance. Agreement and discrimination improved for high-confidence fine-mapped variants, and Borzoi and AlphaGenome showed stronger inter-model convergence on fine-mapped variants than on nominal eQTLs, consistent with enrichment for regulatory variants whose effects are more apparent to sequence-based models. In distance-matched AUROC analyses at PIP [&ge;] 0.9 using PIP < 0.01 variants as low-PIP comparison variants, the AA high-PIP variant set yielded the highest discrimination for both Borzoi (0.837 [95% CI: 0.790-0.870]) and AlphaGenome (0.820 [0.793-0.845]). The CH-versus-NHW ordering was model-dependent: Borzoi yielded higher AUROC in NHW than CH, whereas AlphaGenome produced nearly identical CH and NHW estimates. AUROC values were lower when intermediate-PIP variants were used as comparison variants, but the AA set retained the highest discrimination. MAF-stratified sensitivity analyses attenuated some ancestry contrasts but did not eliminate the higher AA discrimination pattern. Functional annotation analysis showed that AA high-PIP variants more often overlapped chromatin accessibility and chromatin-contact annotations than NHW variants, despite lower overlap with prior eQTL and sQTL annotation catalogs. ConclusionsBorzoi and AlphaGenome showed limited agreement with nominal eQTL effect sizes, but better distinguished high-confidence fine-mapped eQTLs from low-PIP variants. These results support using S2F scores as prioritization evidence for fine-mapped regulatory variants, especially promoter-proximal high-PIP variants, rather than as standalone predictors of eQTL effect size. The strongest discrimination was observed for the AA high-PIP variant set. Overall, the AA result is best interpreted as stronger separation of high-PIP variants from lower-PIP comparison variants, shaped by fine-mapping resolution, LD, the choice of comparison variants, and annotation composition.

bioinformatics↗

Multi-ancestry Transcriptome-Wide Association Study Reveals Shared and Population-Specific Genetic Effects in Alzheimer's Disease

Alzheimers disease (AD) risk differs across ancestral populations, yet most genetic studies have focused on non-Hispanic White (NHW) cohorts. We conducted a multi-population transcriptome-wide association study (TWAS) using whole-blood RNA-seq and genotype data from NHW (n=235), African American (AA; n=224), and Hispanic (HISP; n=292) MAGENTA participants. Using SuShiE for multi-population cis-eQTL fine-mapping, we identified credible sets for 8,748 genes, improving fine-mapping precision relative to analyses using fewer populations. cis-eQTL effects were largely shared across populations, with a subset showing population-specific regulation. We performed population-stratified TWAS of AD and inverse variance-weighted meta-analysis, followed by gene-level TWAS fine-mapping (MA-FOCUS), prioritizing nine genes (FDR<0.05, PIP>0.8), including established AD loci (BIN1, PTK2B, DMPK) with broadly consistent effects across populations. At BIN1, fine-mapped cis-eQTL variants used in the TWAS prediction model highlighted rs11682128, which is only modestly correlated with the GWAS index SNP rs6733839 (r2 {approx} 0.34), demonstrating how integrating eQTL fine-mapping with TWAS can refine signals beyond sentinel GWAS variants. We also identified an association between COG4 expression and AD in NHW, implicating Golgi-related pathways. Using independent SuShiE-derived models from TOPMed MESA (PBMC), several signals replicated directionally across ancestries, with the strongest statistical support in NHW. Overall, multi-population eQTL fine-mapping improves model interpretability and helps resolve shared and population-specific regulatory mechanisms relevant to AD.

genomics↗

Methylation Clocks Do Not Predict Age or Alzheimer's Disease Risk Across Genetically Admixed Individuals

Epigenetic aging clocks based on DNA methylation patterns across the genome have emerged as a potential biomarker for risk of age-related diseases, like Alzheimers disease (AD), and environmental and social stressors. However, methylation clocks have not been comprehensively validated in genetically diverse individuals. Here we evaluate a set of first-, second-, and third-generation methylation clocks in 621 AD patients and matched controls from African American, Hispanic, and White cohorts. The clocks are less accurate at predicting age in genetically admixed cohorts compared to the White cohort, especially for those with substantial African ancestry. This decreased accuracy holds in >2,500 individuals of European and African ancestry from three additional datasets. The clocks also fail to consistently identify age acceleration in admixed AD cases compared to controls. To explore potential causes for the lack of generalization of the clocks, we intersected clock CpGs with methylation, germline genetic variants, and methylation QTL (meQTL) data from global populations. We find differential methylation between African and European ancestry individuals is common for clock CpGs. Genetic variants rarely disrupt clock CpGs between populations, but a substantial fraction of clock CpGs have meQTL with significantly higher frequencies in African genetic ancestries. Our results demonstrate that methylation clocks often fail to predict age and AD risk when applied across populations and suggest avenues for improving their portability by considering differences in genetic and epigenetic patterns across human populations.

genomics↗